Site Reliability Engineering
SLOs, error budgets, incident response and blameless postmortems — reliability as an engineering practice, not a hope.
// uptime is a feature
Hey, I'm Devesh Saini 👋
Open to full-time SRE / DevOps roles · view CV →
I design, automate and run production infrastructure — Kubernetes, AWS, Terraform, CI/CD and observability. When I'm off-call, I build small mobile apps for fun.
// about

I'm a site reliability engineer who treats production like a product: measured, automated, and boring in the best way. Over the years I've moved workloads to Kubernetes, rebuilt CI/CD pipelines, and put observability in place before incidents demanded it.
Off-call, I ship small mobile apps — partly for fun, partly because building the whole thing keeps my empathy for developers sharp.
// services
SLOs, error budgets, incident response and blameless postmortems — reliability as an engineering practice, not a hope.
AWS and GCP design, multi-region setups, migrations and cost optimization that pays for itself.
Cluster design, autoscaling, hardening and the boring operational discipline that keeps workloads healthy.
Pipelines, GitOps and release engineering — from commit to production without human toil.
Metrics, logs and traces with Prometheus, Grafana and OpenTelemetry. See problems before customers do.
Terraform and Ansible for infrastructure that is reviewed, versioned and reproducible.
// career
Bridge engineering and customers for a leading work-management platform — designing technical solutions, building proof-of-concepts, and translating reliability and integration requirements into working implementations. The engineer in the room when the answer has to actually run in production.
Site reliability for enterprise data pipelines — kept large-scale ETL and analytics workloads healthy for global clients. Built Grafana dashboards and alerting that caught pipeline failures before business reports broke, and automated recurring incident toil across the platforms I supported.
Ran production Java web applications on Linux — deployments, monitoring, incident response, and root-cause fixes. Automated routine operations with scripting and hardened the environments I supported, keeping the applications I owned stable and boring in the best way.
// selected work
// things i've built
iOSA mindful digital mala for naam jap and mantra jap — count, track, and reflect on your daily practice.
iOSOn-call schedules that hand themselves off. Live countdowns, auto-built rotations, and reminders before every handoff.
iOSThe offline DevOps toolkit — cron, CIDR, JWT, YAML and cheat sheets, one tap away.
// live from github
Python-Based Log Analysis & Auto-Remediation Tool
// git journal
Single-node to 3-node HA. Etcd snapshots to S3-compatible MinIO. Longhorn for storage.
They must agree or you drop requests on deploys. Wrote a preStop sleep hook into our base chart.
Handoff notifications now respect quiet hours. First feature built entirely from user feedback.
1,200 rps on a $5 VPS thanks to SQLite + one Node process. Boring tech wins again.
// education & certs
Cloud Native Computing Foundation · Score 94/100
Amazon Web Services
HashiCorp
Quantum University
// author

Self-published
Turn ideas into powerful mobile applications with confidence. The Art of App Development is a practical guide to designing, developing, deploying, and maintaining modern mobile apps. Learn UI/UX design, native and cross-platform development, backend architecture, APIs, app monetization, publishing, and emerging technologies like AI, AR/VR, and IoT—all in one comprehensive resource.

Self
Deliver reliable services, exceed customer expectations, and build a resilient business with confidence. Service Level Mastery is a practical guide to designing, implementing, and optimizing Service Level Agreements (SLAs), Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets. Learn proven strategies to improve reliability, reduce downtime, enhance customer satisfaction, and create a culture of operational excellence that drives long-term success.

Self-published
Master Linux from A to Z with the ultimate command reference. Whether you're a beginner exploring the Linux terminal or an experienced system administrator, this practical guide covers hundreds of essential Linux commands with clear explanations, real-world examples, and best practices to boost your productivity and confidence.
// cooperation
“Devesh rebuilt our deployment pipeline and our on-call stopped being a nightmare within a month. The postmortem culture he introduced outlasted his contract.”
“Rare mix: deep infrastructure knowledge and the patience to teach it. Our cloud bill dropped 40% and the team understood why.”
“Handed him a fragile legacy stack; got back Terraform, runbooks, and a team that ships daily.”
// writing
#AWS, Cloud Computing, Tutorials5 minNew to AWS? This beginner-friendly guide breaks down what Amazon Web Services is, the core services like EC2, S3, and Lambda, why businesses choose the cloud, and how to launch your first project free — no jargon, no headaches.
#SRE6 minMost postmortems produce a document nobody reads. Here is the structure that produces fixes instead.
Most clusters are either wasting half their compute or one traffic spike away from OOMKills. A measurement-first method for sizing requests and limits.
// faq
Cloud migrations, Kubernetes setups, CI/CD pipelines, observability rollouts and reliability audits. Anything from a one-week infra audit to a multi-month migration.
Yes — I work with teams worldwide, async-first, with overlap hours agreed up front.
We start with a short call, I scope the work into a fixed plan with milestones, and every project ends with documentation, runbooks and a handover session for your team.
Yes. The SRE Retainer covers ongoing reliability work, incident support and monthly reviews — or 30 days of support is included with every project.
An SRE keeps production systems fast, available and cost-efficient by treating operations as an engineering problem: defining SLOs, automating away manual toil, building monitoring and alerting that catches issues before users do, and running structured incident response and postmortems when things break.
Both. For early-stage teams the work is usually foundations — CI/CD, infrastructure as code, monitoring and a sane on-call setup done right the first time. For larger teams it is usually reliability at scale: SLOs, incident process, Kubernetes and cloud cost work.
Yes — in a specific order: first remove idle resources (typically 15–25% of the bill, zero risk), then rightsize from real utilisation data, then schedule non-production environments, and only then lock in savings plans. Reliability is checked at every step; savings that create outages are not savings.
Yes. I work remotely with teams across time zones, keep an overlap window for meetings and pairing, and put everything else in writing — decision docs, runbooks and recorded walkthroughs — so the work never blocks on my calendar.
The explicit goal is the opposite of dependency. Every engagement ends with documentation, runbooks and a handover session so your team owns and operates everything I built. You should not need me for day-two operations — only for the next project.
Day to day: AWS, Kubernetes, Terraform, Docker, GitHub Actions and other CI systems, Prometheus and Grafana for observability, and Python or Go for tooling. The principles — SLOs, automation, capacity planning — transfer to whatever stack you already run.
// contact