Blog

AI SRE Agent: How We’re Rethinking Incident Investigation

Picture of Kamen Tarlov
Kamen Tarlov
CEO & DevOps Lead
06.10.2026
Reading time: 6 mins.
Last Updated: 06.10.2026

Table of Contents

It’s 3am. An alert fires.

You open your laptop, and the next twenty minutes look like every other incident: cross-referencing a dashboard, tailing logs in a second tab, trying to remember whether this happened before and what fixed it last time.

By the time you’ve reconstructed the timeline, you’re already most of the way to the fix. The investigation was the slow part, not the decision.

That twenty minutes is where we’re spending our engineering effort right now.

ai sre

Before this was a product, it was our day job.

ITGix’s SRE and DevOps teams have designed, run, and fixed production infrastructure for many clients – from Kubernetes clusters and Kafka pipelines to PostgreSQL fleets, across both cloud and on-prem environments. We’ve been the people getting paged at 3 a.m.

That experience has a shape, and it’s rarely as simple as “check the logs.”

It’s knowing that Kafka consumer lag immediately after a deployment often means a rebalance storm, not a slow consumer. It’s knowing that an OOMKilled JVM pod with a flat heap graph can point to native memory rather than the heap. It’s knowing that when PostgreSQL P99 latency climbs while CPU stays flat, you check lock waits and connection saturation before rewriting a single query.

Every senior engineer carries hundreds of these patterns.

The problem is that most of them live in people’s heads, and a general-purpose model doesn’t have them. It knows what Kafka is. It doesn’t necessarily know what tends to go wrong with it at 3 a.m.

So we’re writing that experience down in two forms the agent can use directly.

For every technology we monitor, we ship runbooks written by ITGix and attached directly to the alert rules.

Each SOP covers what the alert means, what to check first, what usually fixes the problem, and what not to touch.

When an alert fires, the RCA agent pulls in the matching SOP and includes it in the investigation trail, so you can see which part of the playbook it followed.

Your team can also add its own SOPs for your services, and the agent can use both.

An SOP tells the agent what to check.

A skill encodes how an experienced engineer checks it: the specific queries to run, which hypotheses to test first, the thresholds that matter, and the false leads to rule out.

We’re building these skills one domain at a time, starting with incidents our own teams have handled.

A model vendor can provide the model, but it can’t provide this operational context.

What makes the difference isn’t simply a smarter model. It’s the engineering judgment that tells the model where to look, what to test, and what evidence matters.

And because SOPs and skills are plain, reviewable documents rather than knowledge hidden inside a model, you can see exactly what the agent knows, correct it, and add your own.

Most teams aren’t blind. They have dashboards, alerts, logs, and traces.

The gap is that every incident still starts from zero.

The platform knows an alert fired. It doesn’t know that the same failure pattern happened four months ago, what the on-call engineer checked first, or what actually fixed it. That knowledge lives in someone’s memory or in a Slack thread nobody will find again.

That’s where we see a practical role for an AI SRE agent.

Its job is narrow and specific: do the first twenty minutes of investigation before a human has to.

Not replace the judgment call. Do the legwork that judgment call depends on.

This isn’t a speculative bet.

Every major observability and cloud vendor has shipped or is shipping something in this space, and the reported results are consistent enough to take seriously: Microsoft’s Azure SRE Agent has collected diagnostics on 18,000+ incidents and reports saving over 10,000 engineer hours. Datadog’s Bits AI SRE reports up to a 95% reduction in time-to-resolution by testing hypotheses against live telemetry instead of dumping raw data at an engineer. PagerDuty built theirs around persistent memory specifically because remembering what worked last time compounds in value with every incident it handles.

There is an important caveat, though – and one we take seriously.

A rigorous 2025 evaluation by ClickHouse found that today’s LLMs, used naively, do not yet outperform expert SREs at diagnosis. Human SREs still win on raw accuracy.

Where AI clearly does win is speed of synthesis: turning noisy logs and metrics into a coherent timeline, drafting the postmortem, and suggesting where to look first.

That’s the honest scope of what this technology is good at right now. And it’s exactly the scope we designed around: an AI agent that accelerates incident investigation and drafts the analysis, with a human still making the call.

Not an autonomous replacement for one.

Three decisions shape everything we’ve built.

Instead of letting the model wander – call a tool, see what comes back, call another tool, repeat until it gives up or gets lucky – our agent forms explicit hypotheses first, investigates only what’s relevant to each one, and shows its reasoning.

You can see why it landed on a conclusion, not just the conclusion.

Not every alert deserves full analysis.

Before anything reaches the RCA agent, it passes through a scoring pipeline grounded in published techniques – entropy scoring, Bayesian signal fusion, survival analysis for escalation speed – that decides how urgently (or whether) it’s worth attention.

It’s tunable, and you can simulate a config change against your own alert history before turning it on. That means you don’t have to guess whether a change will actually help.

The agent reads your actual Prometheus metrics, your actual SOPs, and your actual alert history – not a generic model of what a typical environment looks like.

Onboarding means connecting what you already have, not replacing it.

The agent can investigate freely.

Anything that acts – silencing a noisy rule, restarting a workload, adjusting a threshold – produces a proposal with a preview, and a person approves it before it executes.

Full audit trail, every time.

The handful of exceptions to that rule – regulatory-latency actions where hesitation itself is the risk – are deliberate, documented, and rare. They are not the default.

The current capabilities focus on investigation, analysis, and recommendations:

  • Ask questions in plain language. “How many critical alerts did we have last week on production?” “How was this alert resolved last time?” – answered from your real event history, not a canned response.
  • Root cause analysis grounded in your runbooks. When an alert fires, the agent investigates using your actual SOPs and environment context, not a generic template.
  • Rule and threshold suggestions backed by your real metric history. Not a best-practice guess – a statistical read of your last 30 days, showing you exactly which alerts are under-tuned or over-tuned and why.

We’re actively building toward the agent being able to act, not just investigate – proposing a fix, showing a preview, and executing once you approve it.

Alongside that, we’re working on:

  • a cost-optimization agent that tracks whether its own savings recommendations actually paid off;
  • a security-posture agent built on the same trust model as the SRE agent; and
  • long-term memory, so the next incident benefits from every incident before it, including related tickets already sitting in your Jira.

Keep an eye on the ITGix blogs for the next chapter.

Newsletter for Tech experts

Signal, not noise -

straight to your inbox.

Join 12,000+ engineers and business leaders getting field notes on SRE, DevOps and cloud- native reliability.

Deep-dive tech blogs & case studies
Emerging tech, curated

Your Work Email

We respect your inbox. Read our Privecy Policy

More Posts

When working with Terraform, the default meta-argument for creating resources should almost always be for_each – especially for infrastructure that is expected to live for a long time and evolve...
Reading
Managing secrets securely in Kubernetes is a critical challenge for modern cloud-native environments. Application credentials, certificates, private keys, and passwords must be handled in a way that is secure, auditable,...
Reading
Get In Touch
ITGix provides you with expert consultancy and tailored DevOps services to accelerate your business growth.
Newsletter for
Tech Experts
Join 12,000+ business leaders and engineers who receive blogs, e-Books, and case studies on emerging technology.