Know Your Stack Before You Add AI: StackGen's Lead SRE on Where to Start
The AI SRE Files is our series where we sit in on conversations with practitioners about AI in the SRE world and pull out one idea worth keeping.
This piece draws from "From Tribal Knowledge to Agent Context: How SREs Become Knowledge Engineers," a deep-dive panel at AISRENext Bengaluru held on June 12, 2026.
Kunal Dabir, VP Engineering at StackGen, framed the exercise with a practical constraint: nobody gets unlimited time to review a team's knowledge layer. Given two hours, where would you start?
Aakash Dabrase, Lead SRE at StackGen, wouldn't start with the knowledge layer.
The Foundation Question Comes Before the AI Question
AI readiness starts with the systems behind the knowledge layer. Before evaluating documentation or prompts, verify that engineers can already access service ownership, deployment history, source code, and operational context through the tools they use every day. The knowledge layer becomes valuable only after those fundamentals are in place.
"I wouldn't even go to the AI part. The foundation has to be really strong for AI to be adopted, whether it's for SRE, coding, or anything else you're trying to do."
An Alert Is Only as Useful as Its Metadata
"Let's say you've set up your observability. If I get an alert, do I also get the GitHub repo link for the microservice that the alert belongs to? If that's not the case right now, you really have to go back and start fixing your observability metadata."
Aakash’s point starts with metadata, but the same problem extends to the telemetry itself. In another AI SRE Files conversation, banking technology leader Lalit Mittal explains why reliable AI starts with reliable telemetry, and what happens when the data underneath it is noisy, fragmented, or outdated.
Documentation Shouldn't Replace System Context
The instinct to solve every gap by writing more documentation misses a larger problem. Organizations with well-structured operational data have a head start on AI adoption because agents can retrieve information directly from the systems where it already exists.
Documentation still has an important role to play as context and decision-making guidance. It can't compensate for missing integrations or operational metadata
What a Helm Chart and an AI Copilot Can't Fix on Their Own
A proof of concept can succeed with minimal instrumentation. Production systems place a different set of demands on observability.
"If you just took an observability platform, downloaded a Helm chart, put it on your infrastructure, and now you're going to use AI to help with RCA or incident response, it might work. But I don't think you're going to have any success taking it to scale."
That gap between a proof of concept and production comes down to telemetry completeness, integration reliability, and consistent metadata. Those capabilities determine how well AI performs once the initial pilot gives way to day-to-day operations.
The Integration Audit
The quickest way to assess AI readiness is to start with your existing integrations.
"I'd look at all the integrations currently set up, without AI in the picture, and see what information I can get out of them. That gives me a good idea of whether you should even be doing this in the first place."
A strong integration layer already exposes service ownership, source code, deployments, dashboards, and other operational context. The gaps you find there are the same gaps an AI agent will encounter.
There’s another reason to make operational context accessible through systems instead of relying on what people remember. Pocket FM’s Abhishek Kundalia shares why incident investigations increasingly start with what the system knows.
A Better Starting Point for AI Readiness
Aakash's version of the audit is simple enough to run this week: pull a recent alert, trace the context it provides automatically, and see how far that gets you before anyone has to explain the system to an AI.
The exercise reveals how much operational context your existing systems already provide. The stronger that foundation, the easier it becomes for AI agents to reason about your environment. Any gaps you uncover become the roadmap for improving AI readiness.
Before investing in AI, it’s worth understanding how much operational context your systems already provide. The gaps you find there are a good indication of where the work needs to happen first. Aiden for SRE Community Edition is free to try for teams ready to put that foundation to use.
About StackGen:
StackGen is the pioneer in Autonomous Infrastructure Platform (AIP) technology, helping enterprises transition from manual Infrastructure-as-Code (IaC) management to fully autonomous operations. Founded by infrastructure automation experts and headquartered in the San Francisco Bay Area, StackGen serves leading companies across technology, financial services, manufacturing, and entertainment industries.