AI systems · Automation · Platform engineering

Systems that work. Even when reality doesn’t.

I design and ship production AI systems, business automation and cloud platforms, with the observability, failure handling and handover discipline of fifteen years in production.

A schematic of one client system, not a live readout. What the figures in it are, and what the system does, are set out in the case study.

15+ years in production systems Former VP of DevOps 24/7 broadcast and live-streaming platforms Hands-on delivery English · Hebrew · Russian

One system, as it is built

The interesting part of an automation is what happens when it fails.

A professional services firm runs client intake, documents, tasks and customer updates through one tracked process. It cleared 25,007 runs in ten days and one of them failed. What follows is the part that keeps it running.

PRODUCTION SYSTEM SNAPSHOT ONE SELECTED SYSTEM
  • 01 20 live workflows of 30 built, one tracked process BUILT
  • 02 12 external systems behind it BUILT
  • 03 Centralised failure handling, one alert to one person BUILT
  • 04 Cross-channel deduplication before anything is filed BUILT
Read the case study

What I do

Modern AI and automation engineering, backed by production infrastructure work.

The combination is the point. Plenty of people can call a model or draw a workflow. Fewer have been the person paged when it stops at three in the morning.

01

AI Systems & Agents

Agents that take multi-step action, and production AI systems around them. The model proposes. A check decides.

  • Agentic workflows
  • Tool use / function calling
  • RAG / knowledge systems
  • Document & image understanding
  • Structured outputs
  • Validation gates
  • Guardrails
  • On-device / local inference

Selected stackOpenAI · Gemini · MCP · Claude Code · Codex · Cursor

02

Automation & Integration Engineering

Systems, people and data connected into workflows that survive real operational inputs.

  • Workflow architecture
  • Idempotency & deduplication
  • Retries & backoff
  • Quota-safe writes
  • Failure detection & alerting
  • Backup & recovery
  • Human-in-the-loop
  • Data migration & backfill

Selected stackn8n · REST APIs · Webhooks · Slack · Trello · Google Workspace · WhatsApp

03

Platform, Cloud & Reliability

Production infrastructure, delivery and observability built for failure, recovery and clear ownership.

  • Infrastructure as Code
  • Container orchestration
  • Cloud & hybrid architecture
  • CI/CD & release engineering
  • Observability
  • Incident response
  • Disaster recovery
  • 24/7 production operations

Selected stackAWS · Kubernetes · Terraform / OpenTofu · Docker · GitHub Actions · Prometheus · Grafana · Datadog

04

Software Development & Feature Engineering

Features in software you already run, and the applications and checks around it. Written to be handed back.

  • Feature development on existing systems
  • Taking over inherited systems
  • Custom application development
  • Desktop applications
  • Automated checks in CI
  • Review discipline on AI-written code
  • Handover & documentation

Selected stackTypeScript · React · Python · Rust · Tauri · GitHub Actions

05

Digital Systems & Product Engineering

Websites, internal tools and portals built as part of the business behind them.

  • Information architecture
  • Multilingual & localisation architecture
  • SharePoint solutions
  • Performance & accessibility
  • SEO foundations
  • Owner-controlled hosting & handover

Selected stackSharePoint Online · SPFx · React · TypeScript

Selected work

Three systems, three different disciplines.

All selected work

Built by someone who has carried the pager

“A workflow that only works with clean inputs is a demo.”

Real operations include duplicate events, missing fields, slow APIs and people replying from the wrong address. The recovery path is part of the product.

Dan Burdetsky

Dan Burdetsky
AI Systems, Automation & Platform Engineer
Independent engineer and consultant, 2025 to present

Where the reliability habit came from

Reliability isn’t a feature I add later. It’s the lens I build through.

From 2021 to 2024 I was VP of DevOps for an interactive video and live-streaming platform, through its acquisition. I owned hybrid AWS and on-premises infrastructure, CI/CD, observability, disaster recovery and production reliability while leading a hands-on engineering team.

These were systems used by national broadcasters, where an outage is public, expensive and very hard to hide. The platform was used by broadcasters including NBC News, MSNBC and ABC.

That background shapes how I build now. I expect things to fail, make the failure visible, and design a safe way back.

VP of DevOps, 2021 to 2024 Hybrid AWS and on-premises 24/7 broadcast reliability Hands-on delivery
More about Dan

How engagements work

Advice, the build, or somebody watching it after.

Project-based and advisory engagements. You approve a written scope and a fixed quote before anything is built, and you own every account at the end of it.

01

Architecture & advisory

Review the system you already have, find the real constraint rather than the loudest symptom, and design the approach. Sometimes the answer is that you do not need the build.

02

Build & delivery

Own the implementation from architecture through production handover: the workflows, the code around them, the failure handling, the documentation and the accounts in your name.

03

Support & maintenance

Scoped separately, after handover, for a system that wants somebody watching it. Declining it is a real option and costs you nothing: the handover is complete either way, and nothing in the build depends on me being reachable.

The five practices in full

Notes

Four things I have written, most of which start with something breaking.

All insights

Start with what exists

What are you building, automating or trying to make reliable?

Tell me what exists today, where the hard part is, and what a good outcome looks like.

Message Dan on WhatsApp (opens in a new tab)

or email dan@burdetsky.xyz

No sales funnel. Normally a reply the same business day.
Start a project