Platform & reliability · Experience snapshot
Where the reliability habit came from.
15+ years across infrastructure, production operations and broadcast/streaming platforms, including VP of DevOps. This is where the reliability discipline on the rest of this site came from.
This page is an experience snapshot and not a case study, and the difference matters. The case studies on this site describe systems I designed and built as an independent engineer, and they name what changed. This describes employed roles at other companies. There is no client, no engagement and no deliverable to point at, and the systems belong to the businesses that ran them.
It is here because the discipline in every case study on this site came from it, and hiding that on an About page was the single largest thing wrong with the previous version of this site.
The role
Owning the platform a broadcast product ran on.
From 2021 to 2024 I was VP of DevOps for an interactive video and live-streaming platform, through its acquisition. I owned hybrid AWS and on-premises infrastructure, CI/CD, observability, disaster recovery and production reliability while leading a hands-on engineering team.
These were systems used by national broadcasters, where an outage is public, expensive and very hard to hide. The platform was used by broadcasters including NBC News, MSNBC and ABC.
That background shapes how I build now. I expect things to fail, make the failure visible, and design a safe way back.
What that covered
The parts of it that still show up in the work.
- 01 Hybrid AWS and on-premises infrastructure OWNED
- 02 CI/CD and release automation OWNED
- 03 Observability across production OWNED
- 04 Incident response and root cause analysis OWNED
- 05 Disaster recovery and security hardening OWNED
- 06 24/7 production, broadcast and live-streaming systems OWNED
Open the diagram description
An experience snapshot of employed platform work, drawn as three horizontal bands with a narrow column down the right side. The first band, labelled runs on, shows hybrid infrastructure as two boxes side by side: Cloud, described as capacity that moves with demand, and On premises, described as capacity that has to be planned. A two headed connector joins them, marked hybrid, because the estate spans Amazon Web Services and on-premises hardware at the same time. The second band, labelled gets there by, shows continuous integration, continuous delivery and release engineering as a pipeline of four numbered steps reading left to right: build, then test, then release, then deploy, each joined by an arrow. The third band, labelled is watched by, shows observability feeding an incident path: four small boxes in a row, metrics, logs, alert and on-call, joined by light connectors, and then one bold arrow onward into a highlighted box called root cause analysis. From root cause analysis a long arrow travels upward and back to the left, re-entering the first step of the pipeline, carrying the note that the fix goes back through the pipeline, so the loop closes. The column on the right, outlined in the accent colour, is headed disaster recovery and holds two short lines: a way back, designed at the start, and rehearsed rather than documented. A line along the bottom reads: around the clock production, an outage here is public, and there is no quiet window.
Drawn from this page’s own copy. It is the shape of the role, not any employer’s architecture, and it names no company, product or platform.
Leading it
Five to ten engineers, and about forty before that.
I led a hands-on engineering team of five to ten, and had previously led global technical and support teams of up to about forty, working with teams and clients across Europe, the Americas and Asia Pacific.
Before that: Head of Professional Services at CloudZone by Matrix, Senior Manager for Customer Care at Avid Technology, and Global Support Manager and Tier 3 Team Leader at Orad Hi-Tec Systems. A selected list rather than a complete employment history, and the rest of it is on the About page.
Why it is on a work page
“Built by someone who has carried the pager.”
An automation that fails quietly is the same failure as a stream that fails quietly, and the person who has been woken up by the second one builds the first one differently.
Now
The same discipline, applied on my own account.
Since 2025 I have worked as an independent engineer and consultant. That is a separate chapter rather than a continuation of the role above, and the systems on the rest of this site are from it.
Related service: Platform, Cloud & Reliability
Before it is an incident
What would have to fail before anyone found out?
Tell me what runs in production and what watches it. I will tell you where I would look first, and whether the answer is a tool or an owner.
Ask Dan on WhatsApp (opens in a new tab)or email dan@burdetsky.xyz