A practical cloud cost review for AI teams
A bill review worksheet with a worked example and rollback criteria
Start with the bill and a useful unit
A GPU bill alone cannot tell you whether inference is expensive. A busy server can be good value, and a cheap server can spend most of its time doing no useful work. Start by comparing cost with completed requests, generated tokens, or completed batch jobs. Choose the unit that represents what your customer receives.
This is a review worksheet, not a promised savings percentage. Use your own invoices and traffic measurements. Keep the billing period, currency, region, discounts, and tax treatment consistent when comparing alternatives.
Separate the costs you can change
Break the bill into serving compute, training or experiments, storage, databases, and networking. Give each line an owner and an explanation. A resource without an owner is a reason to investigate, not permission to delete it.
| Check | Evidence to collect | Decision |
|---|---|---|
| Idle endpoints | Request rate, GPU utilization, startup time | Keep warm capacity for interactive traffic; consider scheduled capacity for batch work. |
| Oversized workers | Peak memory, queue age, latency under load | Test a smaller allocation against the same workload. |
| Network charges | Source, destination, bytes, region and billed service | Compare routing alternatives including their fixed and usage costs. |
| Repeated generations | Cacheable request share and correctness constraints | Test caching where a previous response remains valid for that user. |
| Storage and snapshots | Owner, retention policy and restore requirements | Remove only after confirming retention and recovery needs. |
Calculate net savings, not the smallest invoice
Here is an illustrative comparison in USD. It is not a client result or a current provider quote.
| Monthly item | Current setup | Proposed setup |
|---|---|---|
| Serving compute | $6,000 | $4,000 |
| Storage and networking | $500 | $700 |
| Additional operating effort | $0 | $600 |
| Total compared cost | $6,500 | $5,300 |
The compute line falls by one third, but net monthly savings are $1,200, about 18.5%. If migration costs $3,600, simple payback is three months at unchanged traffic. The operating-effort line represents additional effort relative to the current setup, not an assumption that the current system needs no maintenance. Include hardware amortization, spare capacity, support, and licenses where they apply.
Test one change against the same workload
For inference, record the model revision, quantization, input and output lengths, concurrent requests, cache hit rate, and failed requests. Measure time to first token separately from total response time. A shorter answer can look faster while giving the customer less useful output.
For Kubernetes, check pod requests and limits before changing node provisioning. Moving an oversized request onto a cheaper node does not fix the allocation. Keep interruption-sensitive work separate from disposable workers, and rehearse recovery before relying on Spot capacity.
For network changes, trace the actual route first. A private endpoint may help for an eligible destination, but it does not make all network charges disappear. Compare the specific endpoint's pricing and traffic path with the route it replaces.
Finish with an owner and a rollback condition
Write down the change, expected monthly effect, implementation cost, owner, and acceptance criteria. For example: keep the smaller worker only if queue age and error rate remain within your existing service targets at peak load. Record the previous configuration so the decision is reversible.
Review the next comparable billing period after deployment. If traffic changed, show both the total invoice and cost per useful unit. That is the difference between a plausible spreadsheet and an improvement you can explain to the team.