AI Agents are Ops Work

Vault note, not reviewed against the source. Written in the knowledge vault on 2026-04-25 by models working with Stefan Coetzee and published as it stands, with private addresses, e-mail addresses and an employer name redacted. Check claims against the primary source before relying on them.

Managing the lifecycle of an AI agent in production is operational engineering, not developer engineering. Agents need SLOs, error budgets, paging, and runbooks. Treat them like services because they fail like services.

The category mistake

When companies first deploy LLM agents, the agents typically land on a developer team's plate. The result is predictable: outages happen during business hours because nobody is on standby, drift goes unnoticed because nobody is measuring quality, costs explode because nobody owns capacity, and prompt regressions ship because nobody runs canary deployments.

This happens because the developer team that built the agent treats it as their feature. Once the agent is in production, it stops being a feature and starts being a service. Services need operations.

What an agent SLO looks like

Agents have measurable behavior at multiple layers:

LayerExample SLICommon SLO
Latencyp95 time-to-first-token, p95 full response< 3s TTFT, < 30s full
Availabilitysuccessful request rate excluding user errors99.5% over 30 days
Qualityrubric-graded response score on canary set> 0.85 on weekly batch
Costtoken spend per request, per user, per daywithin budget envelope
Tool reliabilitytool-call success rate, tool-call latency99% success, p95 < 1s
Compaction integritysurvives N compactions without behavior drifttested per release

If you cannot answer these for a production agent, you do not have an operations posture. You have wishful thinking.

What pages an SRE for an agent

  • LLM provider outage (cascade if no fallback)
  • Tool dependency outage (MCP server, vector DB, API gateway)
  • Quality regression on canary set (often the silent killer)
  • Cost spike (runaway loops, jailbroken sessions, prompt injection)
  • Context-window saturation as steady-state behavior
  • Auth/secret expiry on tool credentials

These are operational events. They wake someone up. They have rollbacks. They generate postmortems.

The runtime is the new container

Just as containers introduced a new operational layer (Kubernetes, runtime security, container escape) that ops teams had to absorb, AI agents introduce a new operational layer: prompt regressions are the new dependency upgrades, context exhaustion is the new memory leak, tool-call failures are the new network partitions. The pattern repeats. Ops absorbs.

Governance and risk frameworks for agentic systems

See also

Compaction is the New OOM · Memory Architecture L0-L4 · Graph-RAG over Flat RAG for Operational Knowledge · README (manifesto) · 04-incident-management · ISO 42001 Cluster · OWASP LLM Top 10 Cluster