Skip to main content
HB
Top Rated · 100% Upwork JSS6+ Yrs Experience

I keep production up, cut the cloud bill, and run AI infrastructure that doesn't fall over.

Senior SRE & Platform Engineer · FinOps · LLM Infrastructure

Kubernetes clusters that scale themselves. AWS & GCP bills that shrink instead of creep. GPU inference that costs a fraction of the closed-API bill. $160K+ documented cloud savings across 18+ systems.

$2,500 Deep Audit: If I don't find at least $10,000 in annualized cloud/AI waste in 2 weeks, it's 100% refunded.
Delivered remotely for:AUAustraliaCACanadaUSUnited States
FinOps Teardown · Real Client Data

Startup Cloud Spend: 73% Reduction

Annual Savings+$124.8K / yr
Before Optimization
$14,200/mo
Unmanaged SaaS API + Idle GPUs
After Optimization
$3,800/mo
73% Saved · 3-Week Setup
GPU Inference (vLLM on A100 vs SaaS Tokens)$8.4K → $2.1K
Kubernetes Compute (Karpenter Spot vs Fargate)$3.6K → $950
Vector Database (PostgreSQL pgvector vs SaaS)$1.2K → $450
Network Egress (Direct VPC Endpoints vs NAT)$1.0K → $300
p95: 1450ms → 420msSLA: 99.98% MeasuredRead Case →
Remote-First Availability:US Pacific (PST), US Eastern (EST) & Australia (AEDT) Overlap
How I Work →
Production Systems
18+
Measured Uptime
99.9%
Verified Cloud Savings
$160K+
Upwork Job Success
100%
Years Experience
6+ Yrs
Interactive FinOps ROI EstimatorReal-Time Teardown Modeler

Cloud & GPU Infrastructure Savings Calculator

Select your cloud environment and drag the sliders to model exact recoverable waste and preserved runway.

Select Primary Cloud Platform:
Monthly Cloud & AI Spend$35,000 / mo
$5,000
$100,000+
Dedicated GPU Fleet / Compute Nodes4 GPUs / Nodes
0 (CPU only)
16+ GPUs
Current Architecture & BottleneckAWS Profile
Projected FinOps ROIBased on measured metrics across 18+ systems
~65% Cost Cut
Est. Monthly Reduction$17,438Eliminated recurring burn
Annualized Preserved Runway$209,256Direct bottom-line savings
Payback on $2,500 Audit~0.6 weeks
Target Production SLA99.98% Measured
Audit Guarantee$10K+ waste or 100% refund
Direct Calendar BookingZero sales pitch · Engineer-to-engineer

Production Systems I Have Built

Real systems I built and deployed. Each one solved a specific engineering problem. Click any project for the full case study.

MLOps & GPU ServingFreelance client

Self-Hosted LLM Inference on Bare-Metal GPUs

Moved 200 attorneys off $18K/month in API token bills onto self-hosted A100s at 82% lower cost with 420ms p95 latency.

82%Inference Cost Reduction vs. Closed APIs
llama.cppvLLMNVIDIA A100CUDA 12.2+9
Explore GPU Architecture
Platform & SREFreelance client

EKS at Scale: Karpenter Autoscaling & JVM SRE

Rebuilt cluster autoscaling with Karpenter spot-first provisioning — cut compute spend 70% while keeping 99.98% uptime.

<82sScale-Up Latency (Down from 4.5 Minutes)
AWS EKSKubernetesKarpenterHelm+8
Inspect Karpenter Blueprint
Platform & SREBuilt at QisstPay (2022–2024)

Hybrid Cloud Infrastructure for a BNPL Fintech

Unified AWS + GCP infrastructure for a fintech handling 1M+ daily transactions across 17 microservices under PCI DSS audit.

1M+Daily Transactions on Infrastructure I Operated
AWS ECSGCP Cloud RunTerraformMySQL+10
DEMOInspect PCI DSS Infra
AI & RAG PlatformsFreelance client

Multi-Stage Legal RAG with Azure AI Search

Built a 5-layer retrieval pipeline over 3,145 dense legal documents, fixing real failure modes at each step.

0.98Ragas Faithfulness Score (Up from 0.72)
PythonFastAPINext.jsAzure AI Search+8
Review RAG Pipeline
Platform & SREFinTech Client Engagement

Payment API Observability & Latency SLOs

Consolidated fragmented metrics into Datadog APM with strict latency SLOs, cutting recovery time from 40+ min to under 9 min.

78%MTTR Reduction (99.99% Peak Availability)
DatadogAWS LambdaPythonDocker+4
Inspect Architecture
Remote Reliability

How I Work with US, Canada & Australia Engineering Teams

Hiring external infrastructure talent shouldn't mean communication anxiety, missed deadlines, or security risks. Here is how I ensure reliable, senior-level delivery from day one:

Daily Working Overlap on Your Hours

Fear: "Will there be painful timezone delays and laggy responses?"

Dedicated daily working overlap with US Pacific (PST/PDT), US Eastern (EST/EDT), and Australian Eastern (AEDT/AEST). Regular scheduled syncs, live debugging sessions, and zero 24-hour turnaround bottlenecks.

US West (PST): 8 AM – 1 PM Overlap
US East (EST): 9 AM – 3 PM Overlap
Australia (AEDT): Full Working Day Overlap

Engineering Candor & Proactive Tradeoff Modeling

Fear: "Will they just nod and build a brittle system that breaks under load?"

Senior engineering candor. If a proposed architecture has cold-start vulnerabilities, over-provisioned GPU waste, or single points of failure, I flag it in the design phase with concrete tradeoffs and cost models before writing code.

Detailed Architecture Decision Records (ADRs)
Written Tradeoff & Failure Analysis
Loom Video & Async Walkthroughs

SOC 2, GDPR & Enterprise Security Discipline

Fear: "Is company intellectual property, data, and cloud access secure?"

Bank-grade security hygiene: IAM least-privilege policies, short-lived STS credentials, zero plaintext secrets in code, encrypted data in transit and at rest, and automated audit trails compliant with SOC 2, HIPAA, and GDPR standards.

HashiCorp Vault & AWS Secrets Manager
Private VPC Subnets & Bastion Zero-Trust
Audit Logs & Immutability

Self-Documenting Infrastructure — Zero Lock-In

Fear: "Will your team be left with undocumented spaghetti code you cannot maintain?"

Everything is codified in Terraform / Kubernetes manifests with complete Mermaid diagrams, structured READMEs, Grafana telemetry dashboards, and interactive runbooks. Your internal team owns 100% of the repository and IP.

100% Infrastructure as Code (IaC)
Mermaid Architecture Diagrams
Step-by-Step Incident Runbooks

Your 9:00 AM is My 9:00 PM — Dedicated Overlap, Every Single Day

Working overlapping hours with San Francisco, New York, Toronto, Sydney, and Melbourne.

Book Free 20-Min Infra Audit
Hands-On Tooling

Interactive Architecture Labs & Simulators

Hands-on technical artifacts replicating real production challenges: GPU serving bottlenecks, FinOps calculators, and Kubernetes self-healing.

bin/lab --id=incident-response

Kubernetes Incident Response Simulator

Automated detection, triage, and recovery in under 42 seconds

The Problem

The production payment pod hit an out-of-memory crash in a 12-node EKS cluster. Liveness probes failed, Horizontal Pod Autoscaler could not scale because the node pool was at capacity, and over two hundred non-critical alerts created notification fatigue. The on-call engineer had to manually SSH into nodes, identify the failing pod, restart it, and verify recovery. That manual cycle averaged forty-five minutes per incident.

Recovery Time
42sfrom 45m
KubernetesPrometheusGrafana
LAUNCH LAB →
bin/lab --id=finops-optimizer

AWS FinOps Cost Optimizer

Line-by-line cloud cost audit that cut $40,920/year

The Problem

The client AWS bill was $6,400 per month for a mid-sized SaaS platform. Contributing factors included over-provisioned EC2 instances in ap-southeast-1 instead of us-east-1, three NAT Gateways routing S3 and DynamoDB traffic across public endpoints at $0.045/GB, absence of a CDN forcing all API responses through the ALB, an idle db.r5.xlarge RDS instance running overnight, and orphaned storage (three unattached 500GB EBS volumes and two unused Elastic IPs).

Annual Savings
$40,92053% cut
AWSCloudWatchCost Explorer
LAUNCH LAB →
bin/lab --id=cold-start-optimizer

18GB ML Model Cold Start Optimizer

Reduced serverless cold start from 2+ minutes to 8 seconds

The Problem

An 18GB model container on Google Cloud Run required over two minutes to cold start because it downloaded model weights from Google Cloud Storage on every scale-from-zero event. Peak memory hit 28GB against a 32GB container limit, causing 504 Gateway Timeouts for initial requests after idle intervals.

Cold Start
8sfrom 2m+
GCP Cloud RungcsfusePython
LAUNCH LAB →
bin/lab --id=terminal-emulator

Infrastructure CLI Emulator

Web-based terminal simulating kubectl, docker, and terraform commands

The Problem

Onboarding engineers on production clusters created operational risks and configuration drift. Provisioning isolated cloud sandboxes cost $800 monthly per engineer and required ongoing maintenance.

Onboarding
3 Daysfrom 2 weeks
ReactTypeScriptNext.js
LAUNCH LAB →

Testimonials

What clients say after working with me on system infrastructure and MLOps pipelines.

Upwork
Hasan delivered a solid end-to-end build — from scraping to deployment. Communication was excellent throughout, and he handled scope changes gracefully. Would hire again without hesitation.
Mark T.Product Owner, UK Planning PortalData Automation & Supabase Deployment
Upwork
Hasan did an excellent job installing and configuring PostgreSQL and K3s on our bare-metal servers. His Ansible automation was clean and well-documented. He also coached our team on DevSecOps best practices.
James R.Systems Administrator, Enterprise ClientDevSecOps & K3s Deployment
Upwork
Hassan quickly deployed our scikit-learn model on SageMaker with a clean endpoint setup. Very responsive and delivered ahead of schedule. The documentation he provided was thorough.
Chen W.ML Engineer, Tech StartupSageMaker ML Deployment
LinkedIn
I hired Hasan for the deployment of our AWS infrastructure and he exceeded expectations. His Terraform modules were clean, well-documented, and production-ready. Highly recommend for any cloud deployment work.
Haseeb K.Business Development Manager, Partner CompanyAWS Cloud Deployment
Upwork
Hasan was an absolute pleasure to work with. He quickly understood our requirements for the legal research platform and delivered a clean, production-ready backend with LLM integration. His DevOps background made deployment seamless.
Sarah L.CTO, AI.ioAI Backend Development
Upwork
Hasan was very friendly, clear, and knowledgeable about RAG architectures. He set up our entire pipeline with pgvector and LangChain, achieving 95% retrieval accuracy on our domain-specific dataset.
David M.Founder, AI StartupRAG Pipeline Development
Upwork
Hasan did an outstanding job helping me with building our CI/CD pipeline from scratch. He set up GitHub Actions, Docker builds, and deployment automation. Our release cycle went from weekly to daily.
Ahmed K.Engineering Lead, SaaS PlatformCI/CD Pipeline Setup
View all reviews on Upwork
100% Job SuccessTop Rated18 Projects
Hasan Butt — Senior MLOps, SRE & FinOps Consultant

Let's Fix Your Cloud Spend

Available for Contract · US/CA/AU Hours

Whether you need to eliminate runaway GPU compute bills, build a reliable multi-stage RAG pipeline, or stabilize Kubernetes operations, I am ready to help.

Remote-First (Dedicated US PST/EST & Australia AEDT Overlap) · PST/EST/AEDT Overlap
Direct Scheduling:Book Free 20-Min Infra Audit

Guaranteed: No sales pitch · Actionable findings in 20 minutes

Send an Async Inquiry

Prefer text over a call? Drop your current architecture or cloud challenge and I will respond within 4 hours.

NDA & IP Protection standardResponse SLA: < 4 Hours