Articles and practical guides
Notes on model serving, retrieval, Kubernetes, and cloud costs. Project reports and illustrative examples are labelled in each guide.
Cost Optimization & MLOps4 min read
When self-hosted LLM inference is worth the operating cost
Compare model quality, latency, memory, and total operating cost before moving inference to your own GPUs.
Cost modelCompare the same workload
Feb 18, 2025Read Full Article
RAG & AI Architecture4 min read
Fixing legal RAG retrieval with Azure AI Search
A practical guide to legal RAG retrieval: preserve document structure, combine search methods, and evaluate failures.
RetrievalStructure before generation
Feb 24, 2025Read Full Article
FinOps & Cloud Architecture3 min read
A practical cloud cost review for AI teams
Review GPU, storage, and network spending with a practical worksheet and an illustrative net-savings calculation.
WorksheetUse your own invoices
Feb 25, 2025Read Full Article
MLOps & Leadership3 min read
How to give an AI cloud budget an owner
Build a cloud budget around workload owners, useful units, and explicit responses to cost changes.
PlanningOwners, costs, and decisions
Mar 1, 2025Read Full Article
AWS & Kubernetes FinOps3 min read
Karpenter on EKS: sizing, scheduling, and cost checks
Review pod sizing, Spot scheduling, readiness time, and billing assumptions before changing EKS autoscaling.
Karpenter v1Example configuration
Jun 15, 2025Read Full Article