AI systems · Automation · Platform engineering
Systems that work. Even when reality doesn’t.
I design and ship production AI systems, business automation and cloud platforms, with the observability, failure handling and handover discipline of fifteen years in production.
A schematic of one client system, not a live readout. What the figures in it are, and what the system does, are set out in the case study.
One system, as it is built
The interesting part of an automation is what happens when it fails.
A professional services firm runs client intake, documents, tasks and customer updates through one tracked process. It cleared 25,007 runs in ten days and one of them failed. What follows is the part that keeps it running.
- 01 20 live workflows of 30 built, one tracked process BUILT
- 02 12 external systems behind it BUILT
- 03 Centralised failure handling, one alert to one person BUILT
- 04 Cross-channel deduplication before anything is filed BUILT
What I do
Modern AI and automation engineering, backed by production infrastructure work.
The combination is the point. Plenty of people can call a model or draw a workflow. Fewer have been the person paged when it stops at three in the morning.
AI Systems & Agents
Agents that take multi-step action, and production AI systems around them. The model proposes. A check decides.
- Agentic workflows
- Tool use / function calling
- RAG / knowledge systems
- Document & image understanding
- Structured outputs
- Validation gates
- Guardrails
- On-device / local inference
Selected stackOpenAI · Gemini · MCP · Claude Code · Codex · Cursor
Automation & Integration Engineering
Systems, people and data connected into workflows that survive real operational inputs.
- Workflow architecture
- Idempotency & deduplication
- Retries & backoff
- Quota-safe writes
- Failure detection & alerting
- Backup & recovery
- Human-in-the-loop
- Data migration & backfill
Selected stackn8n · REST APIs · Webhooks · Slack · Trello · Google Workspace · WhatsApp
Platform, Cloud & Reliability
Production infrastructure, delivery and observability built for failure, recovery and clear ownership.
- Infrastructure as Code
- Container orchestration
- Cloud & hybrid architecture
- CI/CD & release engineering
- Observability
- Incident response
- Disaster recovery
- 24/7 production operations
Selected stackAWS · Kubernetes · Terraform / OpenTofu · Docker · GitHub Actions · Prometheus · Grafana · Datadog
Software Development & Feature Engineering
Features in software you already run, and the applications and checks around it. Written to be handed back.
- Feature development on existing systems
- Taking over inherited systems
- Custom application development
- Desktop applications
- Automated checks in CI
- Review discipline on AI-written code
- Handover & documentation
Selected stackTypeScript · React · Python · Rust · Tauri · GitHub Actions
Digital Systems & Product Engineering
Websites, internal tools and portals built as part of the business behind them.
- Information architecture
- Multilingual & localisation architecture
- SharePoint solutions
- Performance & accessibility
- SEO foundations
- Owner-controlled hosting & handover
Selected stackSharePoint Online · SPFx · React · TypeScript
Selected work
Three systems, three different disciplines.
Client onboarding and document automation
Intake, documents, tasks and customer updates were moving between five tools by hand. I rebuilt it as one tracked process, and built the failure handling around it first.
VoiceInbox
A desktop app that transcribes voice notes on the machine they are on. The model runs locally, no audio leaves the device, and the one API key it holds lives in the system keychain.
Platform and reliability
Hybrid AWS and on-premises infrastructure, delivery and observability for 24/7 broadcast and live-streaming systems. An experience snapshot rather than a case study, and it says so.
Built by someone who has carried the pager
“A workflow that only works with clean inputs is a demo.”
Real operations include duplicate events, missing fields, slow APIs and people replying from the wrong address. The recovery path is part of the product.
Dan Burdetsky
AI Systems, Automation & Platform Engineer
Independent engineer and consultant, 2025 to present
Where the reliability habit came from
Reliability isn’t a feature I add later. It’s the lens I build through.
From 2021 to 2024 I was VP of DevOps for an interactive video and live-streaming platform, through its acquisition. I owned hybrid AWS and on-premises infrastructure, CI/CD, observability, disaster recovery and production reliability while leading a hands-on engineering team.
These were systems used by national broadcasters, where an outage is public, expensive and very hard to hide. The platform was used by broadcasters including NBC News, MSNBC and ABC.
That background shapes how I build now. I expect things to fail, make the failure visible, and design a safe way back.
How engagements work
Advice, the build, or somebody watching it after.
Project-based and advisory engagements. You approve a written scope and a fixed quote before anything is built, and you own every account at the end of it.
Architecture & advisory
Review the system you already have, find the real constraint rather than the loudest symptom, and design the approach. Sometimes the answer is that you do not need the build.
Build & delivery
Own the implementation from architecture through production handover: the workflows, the code around them, the failure handling, the documentation and the accounts in your name.
Support & maintenance
Scoped separately, after handover, for a system that wants somebody watching it. Declining it is a real option and costs you nothing: the handover is complete either way, and nothing in the build depends on me being reachable.
Notes
Four things I have written, most of which start with something breaking.
-
AI-native Delivery
Why AI writes the code and a check decides if it ships
The model is a component, not the system. What decides whether work ships is a gate that has been watched going red.
-
AI Engineering
My best AI automation habits came from DevOps, not AI
Idempotency, one shared error handler, scoped credentials, alerts and a human review path. None of them is an AI habit.
-
Automation Architecture
A workflow that only works with clean inputs is a demo
Duplicate events, missing fields and slow APIs are normal production conditions, not edge cases.
-
Platform & Reliability
The monitor that went quiet
Two of my own systems were down without me knowing, for two days and for seven weeks. Neither failed loudly.
Start with what exists
What are you building, automating or trying to make reliable?
Tell me what exists today, where the hard part is, and what a good outcome looks like.
Message Dan on WhatsApp (opens in a new tab)or email dan@burdetsky.xyz
No sales funnel. Normally a reply the same business day.