AI Cost Visibility & Optimization Understand, allocate & reduce your AI costs - Learn More

Best Practices for Maximizing Spot Savings and Stability

Spot can save you up to 90% compared to On-Demand prices but it’s often perceived as unreliable.

That’s why we partnered with the AWS Spot team to give you an in-depth overview of how and where to use Spot. We’ll also explain how the nOps platform natively works with Spot to make it easy to confidently adopt Spot in your organization.

Stay tuned for:

  • Spot overview and how it works
  • Spot Best Practices and the ideal workloads for Spot
  • Getting started with Spot and balancing compute purchase types
  • How to run workloads on Spot with confidence
  • Real world example of before and after Spot Optimization

Download the full webinar here:

Related Content

Maximizing Spot Savings and Stability: Webinar

Watch Now

Watch Now
Maximizing Spot Savings and Stability: Webinar

What is AWS Spot and how does it work?

AWS Spot Instances offer an opportunity to utilize spare EC2 capacity at discounts of up to 90%. But what exactly is an EC2 instance? AWS provides a diverse range of over 750 EC2 instance types to suit the compute needs of virtually every single workload — from general-purpose to storage-optimized, memory-optimized, or those with specific processing capabilities like ARM-based instances with Graviton processors.

Typically, users can purchase EC2 instances in several ways:

  1. On-Demand: This is the traditional and most flexible option, where users pay by the second without long-term commitments. It can be used even for unpredictable workloads that may spike unexpectedly. However, it is the most expensive option.
  2. Savings Plans: Over time, you may know what you need in terms of compute footprint for an extended period of time. That’s where Savings Plans come in — these are suitable for long-running stable, predictable workloads, offering significant discounts in exchange for 1 or 3-year commitments.
  3. Spot Instances: These are available at a significant discount as they utilize spare AWS capacity. However, AWS can reclaim these instances with just two minutes’ notice. They are best used with fault-tolerant, loosely coupled, stateless workloads architected to handle these interruptions.

A few Spot misconceptions

Firstly, let’s clear up a few of the most common myths we hear about Spot.

FALSE: Spot Instances involve bidding on prices. Bidding was removed in 2017. Today, Spot pricing is determined by current and long-term supply and demand, leading to relatively stable costs that are easier to predict. You no longer have to deal with the spikes and anomalies that sometimes existed when bidding was still in place.

FALSE: Spot Instances are a secondary, old, or bad instance type. In truth, Spot offers the exact same quality as On-Demand. They are a powerful cost optimization tool, especially when integrated with management solutions like nOps to enhance their reliability and usability within your organizational infrastructure.

FALSE: Spot is just for extreme cost optimization. Spot is good for a number of different types of workloads and not necessarily always just used from a cost perspective.

FALSE: Spot can’t be used in production. While Spot interruptions can occur, with appropriate workload qualification, planning and management many AWS customers successfully run Spot in production without any end user impact.

FALSE: Spot is difficult. We’ll discuss some best practices and solutions like nOps that help ensure that Spot is not difficult to implement.

Spot Best Practices & Tips

Let’s start by understanding some best practices that are absolutely key to Spot success.

Why Spot diversification matters

Understanding Spot starts with comprehending what we mean by “spare capacity.” In each AWS region and Availability Zone (AZ), AWS operates various specific instance sizes to maintain cloud elasticity, ensuring it can accommodate incoming requests for EC2. This inevitably results in some level of spare capacity—essentially unused resources.

The availability of Spot capacity can fluctuate, influenced by seasonal spikes or increased demand from particular organizations for particular instance types, affecting how much spare capacity is available and the potential cost savings.

To effectively leverage Spot, it’s important to open up your doors to as many of these different capacity pools as you possibly can. By flexibly employing a variety of instance types, sizes, and distributing your workloads across different AZs and regions, you can ensure you always have a backup plan in case of an interruption.

Interruptions only occur when AWS must reclaim instances for On-Demand or Savings Plans customers. You’ll receive a two-minute warning, typically sufficient to manage the transition or mitigate the impact on your operations by draining tasks or redistributing them to other capacity pools.

Let’s talk about some of the automations and integrations that can help you do this.

Spot integrations for ease of management

Spot has become significantly easier to adopt due to robust integrations with various AWS native services. This includes integrations with services like Amazon EKS, ECS, Auto Scaling, EC2 Fleet, and EMR, all designed to simplify the implementation of diversification strategies previously discussed and help mitigate the impacts of Spot interruptions effectively.

Beyond AWS native services, Spot has also found extensive application within various open-source tools. For instance, whether you’re using a managed service like Amazon EKS or self-managing your Kubernetes cluster, Spot instances can seamlessly integrate into your environment.

You can also use Spot with your own self-managed Kubernetes system without any issues. For example, Jenkins pipelines are excellent places to start integrating Spot. When you initiate a build, you need compute resources, and once completed, you want to shut them down again. Fortunately, Jenkins has a native integration with Spot, making it a seamless process.

Additionally, Spot-ready partners play a crucial role in facilitating the broader adoption of Spot Instances. These partners, including nOps, have collaborated closely with AWS to incorporate all these best practices of diversification, workload qualification, and interruption-handling directly into their platforms. These partnerships enable customers to leverage EC2 Spot confidently, optimizing costs without compromising on performance or availability.

What is nOps Compute Copilot?

nOps Compute Copilot is an intelligent workload provisioner. It continuously manages, scales, and optimizes all of your AWS compute to get you the lowest cost with maximum stability.

In line with AWS recommendations, Compute Copilot was not built on proprietary auto-scaling technology but integrates with your preferred and existing AWS native services.

AWS native services

Native integration is preferred because it means that for your existing workloads, there’s very low overhead to migrate them to Compute Copilot. Copilot will automatically and sensibly apply commitments like Savings Plans and Reserved instances. It will continuously tune and optimize your workloads for you, putting your configurations for ASGs, EKS auto-scaling, ECS auto-scaling, and Batch on autopilot without constant effort from your engineering teams.

A sidenote on the benefits of Karpenter

Many organizations running workloads on EKS or self-managed Kubernetes on AWS are hearing more and more about Karpenter, which recently went GA. Karpenter is the most advanced EKS node provisioning framework currently available on the market today.

One of the biggest problems when you move to a container orchestration platform like Kubernetes is making sure that you can efficiently provision your containers and pack them onto your AWS compute instances. Unlike traditional methods that rely on standard-sized instances, Karpenter selects the most optimally sized node or EC2 instance based on specific workload needs, choosing from a variety of instance types and sizes to ensure optimal sizing.

Additionally, Karpenter enhances container consolidation by actively seeking opportunities to shut down nodes that are no longer optimally sized or utilized. It also features native support for AWS Spot tools, seamlessly integrating with the AWS Spot ecosystem.

At nOps, we’ve purpose-built the Compute Copilot to facilitate easy integration with Karpenter. For organizations currently using Cluster Autoscaler, we’ve successfully assisted numerous customers in transitioning to Karpenter.  Compute Copilot enhances organizational awareness within EKS clusters, tuning them for optimal performance based on commitment inventory insights and trends within the Spot market. Additionally, it automates Karpenter configuration, reducing the need for engineering teams to continuously audit your Karpenter settings.

Optimize your cloud costs automatically with nOps

Commitment optimization is often the biggest lever for cloud savings. nOps helps you save 50-60% automatically, with 5-minute setup and no infrastructure changes required.

Maximize savings. nOps continuously layers commitments in small hourly increments to capture 55%+ savings while minimizing overcommitment risk — even for dynamic workloads

Results-based pricing model: Customers typically save ~20% more by switching to nOps — and you pay only when you get better results.

Free Savings Analysis: Quantify exactly how much more you can save for no work on your part. We optimize, and you get the credit.

You can book a free savings analysis to find out how nOps can help you start saving today!

nOps processes over $4 billion in cloud spend and was recently named #1 in G2's cloud cost management category.

Tags

nOps

nOps

Published Date: September 5, 2024, Spot

Featured Content

Introducing Cursor Integration in nOps

Announcement

Introducing Cursor Integration in nOps

byRick Haggart
Introducing Claude.ai (Enterprise) Integration in nOps

Announcement

Introducing Claude.ai (Enterprise) Integration in nOps

byRick Haggart
Amazon EMR Cost Optimization: How to Cut AWS Big Data Processing Costs by 30% or More

Cost Optimization

Amazon EMR Cost Optimization: How to Cut AWS Big Data Processing Costs by 30% or More

bynOps
AWS Spot Instances: How They Work and When to Use Them

Spot

AWS Spot Instances: How They Work and When to Use Them

bynOps
AWS Cost Visibility, Allocation & Governance

Cloud Management

AWS Cost Visibility, Allocation & Governance

bynOps
AWS Database & Analytics Cost Optimization

Spot

AWS Database & Analytics Cost Optimization

bynOps