I fix infrastructure that has to stay up.
I run a multi-tenant microVM cloud where the product is other people's untrusted code. That's the lens I bring to client work: isolation, blast radius, recovery, and cost — before elegance.
microVM & sandbox platforms
The specialty. If you need to run untrusted or AI-generated code with real isolation, I've built the whole path in production.
- Firecracker orchestration, snapshot/restore boot paths
- Copy-on-write forking, hibernate/wake, scale-to-zero
- Per-tenant network isolation and egress policy
- Capacity modelling and admission control
Cloud architecture & Kubernetes
GCP and AWS, from a design review to a hands-on rebuild. I'll tell you when Kubernetes is the wrong answer.
- GKE/EKS setup, autoscaling, node-pool design
- Multi-region and failover architecture
- Networking, IAM, secrets, least privilege
- Migration planning off (or onto) managed services
DevOps automation & IaC
Infrastructure that reproduces itself, and deploys that roll back cleanly when the build is wrong.
- Terraform modules and state hygiene
- Ansible for config that re-asserts itself
- GitHub Actions pipelines, blue-green releases
- Killing config drift at the source
SRE, observability & cost
Know something broke before your customers tell you — and stop paying for compute nobody is using.
- Prometheus, ClickHouse, alerting that means something
- Incident response, RCAs, runbooks
- Cloud bill forensics and rightsizing
- SLOs, capacity planning, on-call design
Straight pricing, no discovery-call maze.
Rates are in USD. Fixed-price quotes available once scope is clear — I'd rather quote a number than bill you for surprises.
Hourly
Advisory, reviews, focused fixes
- Architecture & design reviews
- Debugging a specific production problem
- Pairing with your team
- Minimum 4-hour block
Monthly retainer
Ongoing partner, predictable cost
- ~40–60 hours/month of focused work
- Roadmap ownership for your infra
- Async-first, with weekly sync
- Priority response on incidents
Project
A defined outcome, quoted upfront
- Platform build or migration
- Cost-reduction engagement
- Observability & alerting rollout
- Fixed price after a short scoping call
What the first two weeks look like.
Scoping call
30 minutes. What's broken, what's the deadline, what have you tried.
Written plan
A short doc: findings, options with tradeoffs, and what I'd do first.
Ship something
Week one ends with a real change in your repo, not a slide deck.
Hand over
Runbooks and docs so your team owns it when I leave.