devesh
.devops
.sre
_
About
Services
// services
⚙
Site Reliability Engineering
→
⚙
Cloud Architecture
→
⚙
Kubernetes & Containers
→
⚙
CI/CD & Automation
→
⚙
Observability
→
All services
→
Work
// work
◇
Zero-downtime Kubernetes migration
Migration · 2026
→
◇
Observability stack rollout
Monitoring · 2025
→
◇
Multi-region failover on AWS
Reliability · 2025
→
◇
Cloud cost optimization program
FinOps · 2024
→
All work
→
Built
// built
⌘
Jaap Jap
A mindful digital mala for naam jap and mantra jap — count,
→
⌘
ShiftBell
On-call schedules that hand themselves off. Live countdowns,
→
⌘
Opskit
The offline DevOps toolkit — cron, CIDR, JWT, YAML and cheat
→
All built
→
Blog
// blog
✎
AWS Cloud Explained: Services, Benefits & How to Get Started
AWS, Cloud Computing, Tutorials · 2026-07-17
→
✎
Blameless postmortems that actually change things
SRE · 2026-06-12
→
✎
Kubernetes Requests and Limits: A Practical Sizing Guide
kubernetes · 2026-05-14
→
✎
Set SLOs before you build dashboards
Observability · 2026-04-20
→
All blog
→
Learn
// courses
▶
Linux for SRE
Beginner · 8 steps · free
→
▶
Python Basics for SRE, DevOps & Cloud
Beginner · 11 steps · free
→
▶
Git Basics for SRE, DevOps & Cloud
Beginner · 10 steps · free
→
▶
Linux Basics for SRE, DevOps & Cloud
Beginner · 11 steps · free
→
// practice
⌨
Labs
</>
Playground
?
Quizzes
>_
Terminal
My learning
→
Contact
Close
Menu
About
Services
Work
Built
Blog
Experience
Journal
Education
Testimonials
FAQ
Publications
GitHub ↗
Contact →
Learn
← home
// selected work
All work
2026
Zero-downtime Kubernetes migration
Migration ·
40+ services moved, 0 minutes downtime
· read case study ↗
EKS
Terraform
ArgoCD
→
2025
Observability stack rollout
Monitoring ·
MTTR down 65%
· read case study ↗
Prometheus
Grafana
OpenTelemetry
→
2025
Multi-region failover on AWS
Reliability ·
RTO cut from 4h to 8 min
· read case study ↗
Route53
Aurora
Terraform
→
2024
Cloud cost optimization program
FinOps ·
-42% monthly AWS spend
· read case study ↗
AWS
Kubecost
Spot
→
2024
CI/CD overhaul for 30-repo org
Automation ·
Deploys 3/week → 20/day
· read case study ↗
GitHub Actions
Docker
Helm
→
2023
On-call & incident process redesign
SRE ·
Pages down 70%, no more alert fatigue
· read case study ↗
PagerDuty
Runbooks
SLOs
→
⌕
esc
↑↓
navigate ·
↵
open