HB
Back to articles
FinOps & Cloud Architecture3 min read

A practical cloud cost review for AI teams

A bill review worksheet with a worked example and rollback criteria

By Hasan Butt · Updated

Start with the bill and a useful unit

A GPU bill alone cannot tell you whether inference is expensive. A busy server can be good value, and a cheap server can spend most of its time doing no useful work. Start by comparing cost with completed requests, generated tokens, or completed batch jobs. Choose the unit that represents what your customer receives.

This is a review worksheet, not a promised savings percentage. Use your own invoices and traffic measurements. Keep the billing period, currency, region, discounts, and tax treatment consistent when comparing alternatives.

Separate the costs you can change

Break the bill into serving compute, training or experiments, storage, databases, and networking. Give each line an owner and an explanation. A resource without an owner is a reason to investigate, not permission to delete it.

CheckEvidence to collectDecision
Idle endpointsRequest rate, GPU utilization, startup timeKeep warm capacity for interactive traffic; consider scheduled capacity for batch work.
Oversized workersPeak memory, queue age, latency under loadTest a smaller allocation against the same workload.
Network chargesSource, destination, bytes, region and billed serviceCompare routing alternatives including their fixed and usage costs.
Repeated generationsCacheable request share and correctness constraintsTest caching where a previous response remains valid for that user.
Storage and snapshotsOwner, retention policy and restore requirementsRemove only after confirming retention and recovery needs.

Calculate net savings, not the smallest invoice

Here is an illustrative comparison in USD. It is not a client result or a current provider quote.

Monthly itemCurrent setupProposed setup
Serving compute$6,000$4,000
Storage and networking$500$700
Additional operating effort$0$600
Total compared cost$6,500$5,300

The compute line falls by one third, but net monthly savings are $1,200, about 18.5%. If migration costs $3,600, simple payback is three months at unchanged traffic. The operating-effort line represents additional effort relative to the current setup, not an assumption that the current system needs no maintenance. Include hardware amortization, spare capacity, support, and licenses where they apply.

Test one change against the same workload

For inference, record the model revision, quantization, input and output lengths, concurrent requests, cache hit rate, and failed requests. Measure time to first token separately from total response time. A shorter answer can look faster while giving the customer less useful output.

For Kubernetes, check pod requests and limits before changing node provisioning. Moving an oversized request onto a cheaper node does not fix the allocation. Keep interruption-sensitive work separate from disposable workers, and rehearse recovery before relying on Spot capacity.

For network changes, trace the actual route first. A private endpoint may help for an eligible destination, but it does not make all network charges disappear. Compare the specific endpoint's pricing and traffic path with the route it replaces.

Finish with an owner and a rollback condition

Write down the change, expected monthly effect, implementation cost, owner, and acceptance criteria. For example: keep the smaller worker only if queue age and error rate remain within your existing service targets at peak load. Record the previous configuration so the decision is reversible.

Review the next comparable billing period after deployment. If traffic changed, show both the total invoice and cost per useful unit. That is the difference between a plausible spreadsheet and an improvement you can explain to the team.

More engineering notes