When self-hosted LLM inference is worth the operating cost
A workload checklist for comparing APIs, vLLM, and dedicated GPUs
Compare the same job before comparing the price
A managed model API and a self-hosted open model are not automatically equivalent. They can differ in answer quality, context limits, throughput, and the work needed to operate them. A cheaper token is only useful if the system still completes the job your customer needs.
My GPU serving project describes a Llama 3.1 70B deployment on four A100 80GB GPUs across two servers, using llama.cpp and GGUF quantization. Its reported monthly comparison is $18,000 for API usage versus $3,240 for colocation and power. That is an 82% reduction in those listed charges; it is not a complete ownership-cost comparison. Hardware purchase, engineering time, and support need their own accounting.
This guide explains how to evaluate a vLLM alternative. The configuration below is a starting point for testing, not the command used in that llama.cpp engagement.
Write down the workload
Record model and tokenizer revisions, precision, input length, output length, request concurrency, and the proportion of repeated prompts. Keep the same evaluation questions when comparing the managed and self-hosted models. Include requests that fail or time out in the results.
| Measurement | What it tells you |
|---|---|
| Time to first token | How long the user waits before a streamed answer begins. |
| Total response time | How long the complete answer takes at a stated output length. |
| Aggregate output tokens per second | How much work the server completes across requests. |
| Queue age and error rate | Whether the advertised throughput hides delays or failed work. |
| Task evaluation | Whether the cheaper model still produces acceptable answers. |
Leave memory for more than weights
Model weights are only part of GPU memory use. Runtime allocations and the key-value cache also need room. Longer contexts and more concurrent sequences increase the cache requirement. Start with a conservative concurrency limit, inspect memory under load, and then raise it.
Quantization can reduce the weight footprint, but the effect on speed and quality depends on the model, hardware, and kernel. Test it on your tasks. A general benchmark score does not prove that a model preserves the legal clauses, code, or numbers your application depends on.
A vLLM test configuration
Use an installed vLLM release compatible with your GPU and driver, and an accessible model supported by that release. This shell example requires MODEL_ID to be set. Pin the model revision and record the vLLM version in your benchmark notes. It has not been load-tested here.
# Set MODEL_ID to a model you have access to and capacity to serve.
vllm serve "${MODEL_ID:?Set MODEL_ID before starting}" \
--tensor-parallel-size 2 \
--max-model-len 8192 \
--gpu-memory-utilization 0.85 \
--max-num-seqs 32 \
--host 127.0.0.1 \
--port 8000
The example assumes two visible GPUs. Keep the endpoint behind an authenticated gateway before exposing it to clients. Tune memory utilization and concurrency using observed queueing and memory behavior rather than copying a claimed optimum.
Decide what may be cached
Start with exact matches where the answer remains valid. Include the model version, prompt template, authorization scope, and relevant source-data version in the cache key. A matching prompt from another tenant must not expose that tenant's answer.
Semantic caching needs a separate correctness evaluation. Similar wording can refer to different dates, amounts, or permissions. Record false matches as failures and define invalidation before relying on a hit-rate improvement.
Calculate ownership cost and recovery capacity
Add hosting, hardware amortization, power, network, storage, monitoring, spare capacity, and operating effort. Compare the same time period and traffic. Model a quiet month as well as peak demand: dedicated capacity can be uneconomical when it sits idle.
Finally, test what happens when a GPU worker or an entire server is unavailable. A deployment that meets latency targets only when every device is healthy needs either more capacity, a degraded-service plan, or a different reliability target. Keep that cost in the comparison.