Skip to main content
HB
Available for new projects
Typical Project: $5K–$25K · Engagements from $100/hr

Hasan Iqbal Butt

MLOps & RAG Platform Engineer

I build the infrastructure that takes AI from prototype to production. Reliable pipelines, real-time retrieval, zero-downtime deploys across 18+ production systems.

Lahore, PakistanRemote — WorldwideReplies within 4 hours
Initializing...
Production Systems
0+
Uptime
0.0%
Infra Cost Cut
$0K+
Job Success
0%
Years Experience
0+

Experience

DevOps & SRE Engineer

AEITCH
Nov 2025 – PresentLahore, PKremote
  • Designed scalable AWS cloud environments and CI/CD pipelines for high-concurrency SaaS platforms, achieving 99.9% uptime
  • Automated IaC solutions using Terraform, AWS CDK, and CloudFormation for consistent staging/production configurations
  • Integrated AI coding assistants (Claude Code, Cursor, Copilot) to accelerate scripting, reducing infrastructure development times by 40%
AWSTerraformCDKGitHub ActionsJenkinsPrometheusGrafanaELK Stack

Lead DevOps Engineer

Nextbridge
May 2025 – Nov 2025Lahore, PKonsite
  • Deployed a modular .NET + Angular school management system on AWS ECS with CI/CD via CodeBuild and CodePipeline
  • Enforced GitOps best practices via GitHub for version-controlled, audit-friendly infrastructure changes
  • Implemented CloudWatch anomaly detection, reducing false alerts and improving on-call rotation efficiency
AWS ECSCodePipelineDockerCloudWatchPython LocustNessusISO 27001

Senior DevOps Engineer

VentureDive
Mar 2024 – Apr 2025Lahore, PKonsite
  • Deployed large-scale monolithic Java application on AWS Elastic Beanstalk with RDS PostgreSQL using modular Terraform
  • Reduced AWS infrastructure cost by 30% by migrating workloads from Singapore to North Virginia region
  • Enhanced Python automation scripts for gigabyte-scale S3 data transfers and cross-environment file processing
AWS Elastic BeanstalkTerraformRDS PostgreSQLS3CloudWatchPythonBash

DevOps Engineer

QisstPay
Aug 2022 – Mar 2024Islamabad, PKonsite
  • Migrated 17+ backend microservices (Python, Java, Go) to GCP Cloud Run, enabling autoscaling and faster builds
  • Ensured PCI DSS compliance across AWS workloads and internal office network with full audit documentation
  • Established secure VPN tunnels with financial and telecom institutions for credibility scoring API integration
GCP Cloud RunAWSTerraformRedisPostgreSQLPCI DSSVPNDocker

Junior DevOps Engineer

Khired Networks
Oct 2020 – Aug 2022Lahore, PKonsite
  • Containerized applications with Docker, running services on physical staging servers and AWS EC2
  • Organized CI/CD pipelines using Bitbucket, Jenkins, and AWS CodeDeploy for automated deployments
  • Set up Graylog centralized logging with Prometheus and Grafana for real-time server monitoring
DockerJenkinsBitbucketAWS EC2GraylogPrometheusGrafanaS3CloudFront

Production Systems I Have Built

Real systems I built and deployed. Each one solved a specific engineering problem. Click any project for the full case study.

Full-StackFreelance client -- live in production

Markaz Idara Tul Islah — An ERP for an Institute That Ran on Paper

An institute running on paper registers, digitalized end-to-end -- bilingual Urdu/English, live with 200+ users.

200+Active Users — Live in Production
Next.jsExpressPostgreSQLRedis+5
DEMO
Platform & CloudBuilt at QisstPay (2022–2024)

QisstPay — Hybrid Cloud Infrastructure for a BNPL Fintech

Hybrid AWS + GCP infrastructure for a BNPL fintech -- 17 microservices, PCI DSS scope, bank VPN integrations.

1M+Daily Transactions on Infrastructure I Operated
AWS ECSGCP Cloud RunTerraformMySQL+10
DEMO
Platform & CloudFreelance client

Production EKS Platform with SRE Observability

An EKS platform with real SRE discipline -- SLOs, Karpenter, and 99.9% measured over six months.

99.9%Uptime Over 6 Months
AWS EKSKubernetesKarpenterHelm+8
Platform & CloudFreelance client

High-Availability Booking Platform — AWS ECS on EC2

Fargate pricing said $8k/month. ECS on EC2 with Reserved Instances said $3.8k. High-availability booking at scale for less.

99.95%Uptime · 9 Months
AWS ECSEC2ECRTerraform+10
DEMO
Platform & CloudFreelance client

GDPR and SOC 2 Compliant Data Pipeline on AWS

GDPR erasure across 12 systems in 72 hours, with Step Functions execution logs as audit evidence.

72hrsRight-to-Erasure SLA Across 12 Systems
AWS GlueAWS KMSS3AWS Macie+8
DEMO
MLOps & Model ServingFreelance client

Custom LLM Serving on Bare Metal GPUs with llama.cpp

200 attorneys, $18k/month in API calls -- moved to self-hosted A100s at 82% lower cost.

82%Cost Reduction vs Cloud GPUs
llama.cppNVIDIA A100CUDA 12.2GGUF+8
RAG & AI AgentsFreelance client

Multi-Stage Legal RAG with Azure AI Search

Legal RAG that respects document structure: 5 retrieval layers over 3,145 docs, each layer fixing a real failure.

3,145 Docs5-Layer Retrieval Pipeline
PythonFastAPINext.jsAzure AI Search+8
RAG & AI AgentsFreelance client

MCP Server for ERP Integration

214 undocumented ERP tables turned into 42 curated MCP tools an agent can actually use safely.

42 Tools90% of Sampled User Questions Answerable via Tools
MCP ProtocolPythonFastAPIPostgreSQL+6
Process

How I Work

Every engagement follows the same five-phase structure. Clear milestones, weekly deliverables, no ambiguity.

01

Discovery

I audit your current infrastructure, identify bottlenecks, and define what needs to change. Free of charge.

02

Architecture

You get a detailed proposal: system diagrams, cost projections, timeline, and tradeoff analysis. No surprises later.

03

Build

Iterative delivery with CI/CD. Working code ships to staging every week, not just at the end of a contract.

04

Deploy

Production rollout with monitoring, alerts, and rollback plans. SLOs defined before anything goes live.

05

Handoff

Runbooks, dashboards, documentation, and a team walkthrough. You own everything. Zero vendor lock-in.

Engineering Standards

What I Build For

These are baseline requirements I hold every system to, documented in the architecture proposal before any code is written.

0%Uptime Target

Reliability Engineering

Every system ships with defined SLOs, SLIs, and error budgets. I set measurable uptime targets and build the alerting to enforce them. No guessing whether production is healthy.

SLASLOSLIError BudgetsIncident Response
0-LayerObservability

Full-Stack Observability

Metrics, logs, and distributed traces from day one. Custom Grafana dashboards, structured logging, and alert routing so you know exactly what happened, when, and why.

PrometheusGrafanaOpenTelemetryPagerDutyELK
0xTraffic Scaling

Scalable Architecture

Auto-scaling groups, horizontal pod autoscaling, and load balancing designed to handle traffic spikes without manual intervention. Tested with load scenarios before launch.

HPAAuto-scalingLoad BalancingCDNCaching
0+Availability Zones

High Availability

Multi-AZ deployments, automated failover, and tested disaster recovery. Every critical path has redundancy. Recovery time objectives are documented and rehearsed.

Multi-AZFailoverDR PlansRTO/RPOBlue-Green
0%Avg. Cost Reduction

Cost Optimization

Right-sized instances, reserved capacity planning, and spot fleet strategies. I build FinOps dashboards so you see exactly where every dollar goes and where to cut.

FinOpsRight-sizingReserved InstancesCost Explorer
0Plaintext Secrets

Security by Default

IAM least-privilege policies, encryption at rest and in transit, network isolation, and secrets management. Compliance-ready infrastructure from the first commit.

IAMKMSVPC IsolationRBACSOC 2 Ready

Live Engineering Demos

Real engineering systems you can interact with. Each demo replicates a complex latency, cost, or scaling problem I solved in production.

bin/demo --id=incident-response

Kubernetes Incident Response Simulator

Automated detection, triage, and recovery in under forty two seconds

The Problem

The production payment pod hit an out-of-memory crash in a twelve-node EKS cluster. Liveness probes were failing, the Horizontal Pod Autoscaler could not scale because the node pool was already at capacity, and alert fatigue from over two hundred non-critical alerts was drowning out the actual P0 signal. The on-call engineer had to manually SSH into nodes, figure out which pod was failing, restart it by hand, and then verify the fix. That whole cycle averaged about forty five minutes per incident.

recovery time
42sfrom 45m
KubernetesPrometheusGrafana+3
LAUNCH DEMO
bin/demo --id=finops-optimizer

AWS FinOps Cost Optimizer

Line-by-line cloud cost audit that cut $40,920/year

The Problem

Client's AWS bill was $6,400/mo for a mid-size SaaS platform. Over-provisioned EC2 instances running in the wrong region (ap-southeast-1 instead of us-east-1), 3 NAT Gateways routing S3 and DynamoDB traffic through public internet at $0.045/GB, no CDN so every API response served directly from ALB, RDS running a db.r5.xlarge at full capacity 24/7 even at 3 AM, and orphaned resources accumulating — 3 unattached 500GB EBS volumes, 2 floating Elastic IPs, and a dead NAT Gateway from a deleted VPC.

annual savings
$40,92053% cut
AWSCloudWatchCost Explorer+5
LAUNCH DEMO
bin/demo --id=cold-start-optimizer

18GB ML Model Cold Start Optimizer

Reduced serverless cold start from 2+ minutes to 8 seconds

The Problem

An 18GB chatbot model deployed on GCP Cloud Run took over 2 minutes to cold-start. The container had to download the full model weights from GCS on every scale-from-zero event. Max memory allocation of 32GB was barely enough — the model plus the Python runtime consumed 28GB at peak. During business hours, cold starts caused 504 Gateway Timeouts for the first users after an idle period.

cold start
8sfrom 2m+
GCP Cloud RungcsfusePython+2
LAUNCH DEMO
bin/demo --id=terminal-emulator

Infrastructure CLI Emulator

Web-based terminal simulating kubectl, docker, and terraform commands

The Problem

Onboarding new DevOps engineers required giving them access to production Kubernetes clusters for hands-on training. This posed security risks — one accidental kubectl delete could take down a service. Sandbox environments were expensive ($800/mo per engineer) and constantly drifted from production configs, making training unrealistic.

onboarding
3 Daysfrom 2 weeks
ReactTypeScriptNext.js+1
LAUNCH DEMO

MLOps and Cloud Engineering Stack

Self-Healing Cluster (K8s)

Click any active pod to terminate it. Watch the replica scheduler spin up replacement tasks instantly.

api-web-7fd-x9Running
api-web-7fd-z2Running
rag-query-5c8-v1Running
rag-query-5c8-k4Running
worker-queue-2b-a8Running
worker-queue-2b-c9Running
replicaset-controller-log
>Deployment up: 6 replicas running.

Vector Search Matcher

Select a natural query. Witness cosine similarity distance indexing fetch matching document embeddings.

rds-replication-runbook.md98.4% match
monitoring-alerts-setup.tf89.2% match
vpc-peering-routing.json74.1% match

CI/CD Pipeline Flow

Trigger git commit lifecycle rollout. Runs linting, image scanning, model tests, and AWS deployment updates.

1
Lint Code
2
Unit Tests
3
Trivy Scan
4
Docker Build
5
ECS Rollout

FinOps Cost Right-Sizing

Adjust cluster node profile capacity level. Observe the immediate drop in under-utilized CPU waste and billing savings.

Over-ProvisionedRight-Sized
ACTIVE PROFILE: m5.2xlarge
8 vCPU32 GB
EST. MONTHLY BILL
$1420
SAVINGS GENERATED
0%
WASTE RATING
High

Testimonials

What clients say after working with me on system infrastructure and MLOps pipelines.

Upwork
Hasan delivered a solid end-to-end build — from scraping to deployment. Communication was excellent throughout, and he handled scope changes gracefully. Would hire again without hesitation.
Mark T.Product Owner, UK Planning PortalData Automation & Supabase Deployment
Upwork
Hasan did an excellent job installing and configuring PostgreSQL and K3s on our bare-metal servers. His Ansible automation was clean and well-documented. He also coached our team on DevSecOps best practices.
James R.Systems Administrator, Enterprise ClientDevSecOps & K3s Deployment
Upwork
Hassan quickly deployed our scikit-learn model on SageMaker with a clean endpoint setup. Very responsive and delivered ahead of schedule. The documentation he provided was thorough.
Chen W.ML Engineer, Tech StartupSageMaker ML Deployment
LinkedIn
I hired Hasan for the deployment of our AWS infrastructure and he exceeded expectations. His Terraform modules were clean, well-documented, and production-ready. Highly recommend for any cloud deployment work.
Haseeb K.Business Development Manager, Partner CompanyAWS Cloud Deployment
Upwork
Hasan was an absolute pleasure to work with. He quickly understood our requirements for the legal research platform and delivered a clean, production-ready backend with LLM integration. His DevOps background made deployment seamless.
Sarah L.CTO, AI.ioAI Backend Development
Upwork
Hasan was very friendly, clear, and knowledgeable about RAG architectures. He set up our entire pipeline with pgvector and LangChain, achieving 95% retrieval accuracy on our domain-specific dataset.
David M.Founder, AI StartupRAG Pipeline Development
Upwork
Hasan did an outstanding job helping me with building our CI/CD pipeline from scratch. He set up GitHub Actions, Docker builds, and deployment automation. Our release cycle went from weekly to daily.
Ahmed K.Engineering Lead, SaaS PlatformCI/CD Pipeline Setup
View all reviews on Upwork
100% Job SuccessTop Rated18 Projects
Hasan Iqbal Butt

Contact

SYSTEMS ONLINE // LIVE

I am available for freelance work and full-time remote roles. Reach out and I will respond within a few hours.

contact@hasanbutt.com
Lahore, PakistanPKT (UTC+5)
Availability: Remote — Worldwide
Availability: Remote — Worldwide · Usually responds within 4 hours