Limited availability — one build per quarter
AI agents that run your operations — not just your demos.
I design and ship production AI agents for incident response, observability, and internal knowledge. Four of them run today inside a Fortune 500 pharmaceutical enterprise — on a multi-application Kubernetes estate, behind an enterprise incident queue. One cuts mean-time-to-acknowledge by roughly half.
30 minutes · video call · no pitch
4
agents live in enterprise production
~50%
MTTA reduction on live incidents
10
production systems shipped
6
MCP servers & integrations built
The short version
Most AI projects die between the demo and the on-call rotation.
A working prototype proves a model can do something once. It says nothing about what happens at 3am when the ticket is ambiguous, the logs are noisy, and the query times out on high cardinality.
That gap is the whole job. Closing it means meeting your systems where they already are — so I write the tool layer myself, including the MCP servers that let an agent reach Grafana, ServiceNow or Confluence reliably rather than approximately. That layer is usually the difference between an agent that demos and one that gets deployed.
Give me the problem statement. I build the system that solves it reliably, in production, where it creates measurable value.
Selected work
Shipped, not sketched.
Three systems in production. Each one shown the way an engineer would want to see it: the problem, the architecture, the guardrail, and what actually changed.
Incident Copilot
A ticket-triggered triage agent that pulls the right resolution procedure, the relevant logs and comparable prior cases the moment a ticket is created — then posts a recommended fix onto the ticket before an engineer has opened it.
Read the case study →
~50%
reduction in mean-time-to-acknowledge
Observability Copilot
Natural-language access to logs, metrics, traces and live endpoint health across a multi-application Kubernetes estate — built on a custom Grafana-stack MCP server I wrote because the one I needed did not exist.
Read the case study →
0
lines of PromQL an engineer has to write
Autonomous SDLC Pipeline
Three cooperating agents — Coder, Reviewer, Deployment — that let one consultant carry a client's entire production computer vision service without hand-running every fix, test and deploy.
Read the case study →
hours → minutes
feature-to-deploy turnaround
Also built
Company Knowledge Base Agent
Enterprise Knowledge Management · Production
DevOps Diagnostics Agent
Cloud Operations / Autonomous SRE · Production
Multimodal Invoice & Document Extractor
Intelligent Document Processing · Production
Content Automation Pipeline
Content Automation / Brand Systems · Production
How I build
Capability is easy. Restraint is the skill.
Most agent projects ask how much autonomy they can grant.
I start from what the system must never be allowed to do.
Human-in-the-loop by default
Every agent I have put into production retrieves, reasons and recommends — a person decides. The incident copilot posts a recommended fix onto the ticket; the SRE acts on it. That boundary is why these systems were allowed anywhere near production in the first place.
Scoped, non-destructive access
My DevOps diagnostic agent has SSH into a production EC2 host with a hard allowlist of read-only commands. It can tell you exactly what is broken and what the fix should be. It cannot apply it. Capability is easy; deciding what an agent must never be able to do is the engineering.
Built on the stack you already run
No rip-and-replace. I wrap what you have — Grafana, Loki, Prometheus, Jaeger, ServiceNow, Confluence, GitHub Actions — behind MCP servers and tool interfaces. When I needed a Grafana-stack MCP server that didn't exist, I wrote one.
Failure modes before features
High-cardinality PromQL that times out. Ambiguous tickets that retrieve the wrong runbook. Log noise that poisons a summary. I design for these first, because they are what actually decides whether an agent is trusted six months in.
Working together
Engagements
Two ways in. Both start with the same 30-minute call about the problem you actually have.
Most engagements
4–8 weeks
Fixed-scope agent build
One agent, scoped to one painful workflow, shipped into your stack and running in production.
Who it is for
Platform, SRE and DevOps teams who know exactly which manual loop is burning their week.
- Discovery: the workflow, the data sources, the failure modes, the success metric
- Architecture with an explicit guardrail and human-in-the-loop design
- Tool/MCP layer over your existing systems — no platform migration
- Evaluation harness so you can prove it works before you trust it
Also available
Ongoing, no minimum
Hourly consulting
Architecture review, unblocking, and a second pair of eyes from someone who has run agents in production.
Who it is for
Teams already building who need to pressure-test a design, fix an agent that works in demo and fails in prod, or decide whether to build at all.
- Agent and RAG architecture review
- Evaluation strategy — how you will actually know it works
- Guardrail, permission and human-in-the-loop design
- Model, cost and latency trade-off decisions
The process
How a build runs
Problem statement
A 30-minute call. You describe the manual loop that hurts. I tell you honestly whether an agent is the right answer — sometimes it is a script and I will say so.
Scope and success metric
We agree on one workflow, one measurable outcome, and what the agent is explicitly not allowed to do. Fixed price from here.
Build against your stack
Tool layer over your existing systems, evaluation harness alongside it. You see working software early, not a slide deck at the end.
Ship and hand over
Deployed in your environment, observable, documented, owned by your team. Measured against the number we agreed in step two.
What you get
A running system with a measured before and after — deployed in your environment, observable, documented, and owned by your team. Not a proof of concept that dies in a branch.
Limited availability
Tell me the problem.
I’ll tell you if an agent is the answer.
A 30-minute call, no pitch. Describe the manual loop that hurts and I’ll give you a straight read on whether this is worth building — including when the honest answer is a script, not an agent.
30 minutes · video call · no pitch