Karpenter on EKS: sizing, scheduling, and cost checks
A v1 NodePool example and the limits of an autoscaling cost comparison
Fix pod sizing before buying cheaper nodes
The EKS project describes two problems that overlapped: slow capacity changes and Java services exhausting container memory. A 512MB heap inside a 512MB container leaves no allowance for thread stacks, direct buffers, and other off-heap allocations. A different autoscaler cannot fix that memory limit.
The project records a 768MB container limit for that heap size, along with startup-probe changes to allow JVM warmup. Treat those values as context for that application, not a sizing rule. Profile your own service and test its startup path under load.
What changes with Karpenter
Karpenter observes unscheduled pods and provisions nodes that meet their scheduling constraints. The pool configuration controls which capacity it may choose. Pod requests, affinity rules, taints, zone constraints, and available instances all affect the result.
Separate workloads that can tolerate interruption from those needing steadier capacity. On-Demand instances can still fail; replication and recovery remain necessary. A PodDisruptionBudget helps with voluntary disruptions but does not prevent a provider from reclaiming Spot capacity.
An example for the v1 API
This is a NodePool example for Karpenter's karpenter.sh/v1 API. It requires a configured controller, permissions, and an EC2NodeClass named default with suitable subnets, security groups, and an AMI. It has not been applied to a live cluster as part of this guide.
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: application-workers
spec:
template:
spec:
nodeClassRef:
group: karpenter.k8s.aws
kind: EC2NodeClass
name: default
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"]
- key: kubernetes.io/arch
operator: In
values: ["amd64"]
- key: kubernetes.io/os
operator: In
values: ["linux"]
limits:
cpu: "64"
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 1m
For a workload that must use On-Demand capacity, add an appropriate scheduling requirement to the pod template. Ensure a compatible pool exists and test the requirement before rollout.
# Fragment inside a Deployment's spec.template
spec:
nodeSelector:
karpenter.sh/capacity-type: on-demand
Keep the financial comparisons separate
The project page reports monthly compute charges of $3,600 before and $950 after its changes, a reduction of about 74% in the listed charges. Its underlying billing export is not included here.
For your own comparison, show node-hours, instance types, purchase options, region, billing dates, and workload volume. Include migration effort and any spare capacity. Compare cost per completed job or request when traffic changes between periods.
Measure readiness, not just instance launch
Record when a pod becomes Pending, when a node is available, when the image is pulled, and when the application passes readiness. Slow image pulls or startup work can dominate even when instance provisioning is fast. Use repeated runs and report the distribution rather than the fastest example.
Rehearse worker termination and queue recovery. Confirm that retries do not duplicate externally visible work, and that remaining replicas can carry traffic during a disruption. Set consolidation policies with the application's warmup and drain behavior in mind.
When to keep the simpler setup
A small, steady workload may not justify another controller and its operational requirements. Compare Karpenter with well-sized managed node groups using the same capacity and reliability assumptions. Keep the option that meets the service target with a cost and maintenance burden your team can explain.