Services · 03

Platform, Cloud & Reliability. Built for failure, recovery and clear ownership.

Production infrastructure, delivery and observability. The parts of a system that nobody notices until the day everybody notices at once.

What I build here

Infrastructure you can rebuild from a repository.

Cloud and hybrid architecture, described as code rather than as a sequence of console clicks somebody remembers. Containerised delivery, a pipeline that gets a change to production without a person holding it, and enough observability that a failure announces itself instead of being reported by a customer.

Then the unglamorous half: incident response that has an owner, root cause analysis that produces a change rather than a document, disaster recovery that has been rehearsed, and hardening that is applied rather than intended.

What it solves

Six situations, and all of them are the same situation.

WHEN THIS IS THE WORK06 PATTERNS
  • 01 Production stopped and a customer told you first SUITED
  • 02 The environment exists but nobody can rebuild it SUITED
  • 03 Deploying is a person following steps rather than a pipeline SUITED
  • 04 There are dashboards, and no alert anyone trusts SUITED
  • 05 The recovery plan has never been run SUITED
  • 06 When it breaks at night, it is unclear whose job it is SUITED

Every one of those is the same question in a different costume: what does this system do when nobody is watching it?

Why this is a service rather than a line on a CV

“Built by someone who has carried the pager.”

Reliability is not a phase at the end of a project. It is a set of decisions taken at the beginning by somebody who has been woken up by the consequences of the other decisions.

The proof

Systems where an outage is public.

From 2021 to 2024 I was VP of DevOps for an interactive video and live-streaming platform, through its acquisition. I owned hybrid AWS and on-premises infrastructure, CI/CD, observability, disaster recovery and production reliability while leading a hands-on engineering team.

These were systems used by national broadcasters, where an outage is public, expensive and very hard to hide. The platform was used by broadcasters including NBC News, MSNBC and ABC.

That is employed experience rather than a client engagement, and this page says so rather than blurring it. The written-up version is on the work index as an experience snapshot.

Hands-on

What I personally do.

Write the infrastructure code and review what somebody else wrote. Build and fix the pipeline. Put the metrics, the logs and the alert in place, and then break something on purpose to find out whether the alert arrives. Lead the incident and write the analysis afterwards. Rehearse the recovery rather than documenting it.

Infrastructure as CodeContainer orchestrationCloud & hybrid architectureCI/CD & release engineeringObservabilityIncident responseRoot cause analysisDisaster recoverySecurity hardening24/7 production operationsNetworking & DNS

Selected stack

AWSKubernetesTerraform / OpenTofuDockerGitHub ActionsGitLab CIPrometheusGrafanaDatadogLinux

Where it meets the rest

An automation that fails quietly is the same failure as a stream that does.

This is not a separate practice bolted onto the others. It is why the automation work has a shared error handler and a silence detector, why the AI work logs every run, and why every build ends with the accounts in somebody else’s name.

Two of my own systems were down without me knowing, for two days and for seven weeks. Neither failed loudly, and the fix was not a better restart.

What next

Scope, then a fixed number.

Most first projects fall between a two-week scoped build and a two-month one. The number is fixed against a written scope after a discovery call, so you approve it before anything is built. What moves it inside that range, in full.

Before it is an incident

Who gets woken up, and with what in their hand?

Tell me what runs in production and what watches it. I will tell you where I would look first, and whether the answer is a tool or an owner.

Ask Dan on WhatsApp (opens in a new tab)

or email dan@burdetsky.xyz

Start a project