Blog

First-Class Citizens: InMobi's Senior Technical Lead on Fixing Stale Runbooks

Written by Anjali | Oct 5, 2026, 10:15:07 AM

The AI SRE Files is our series where we sit in on conversations with practitioners about AI in the SRE world and pull out one idea worth keeping.

This piece draws from “From Tribal Knowledge to Agent Context: How SREs Become Knowledge Engineers,” a panel discussion at AISRENext Bengaluru.

Kunal Dabir named a problem every engineer recognizes, AI or no AI: documentation starts drifting almost as soon as it is written. A GitHub README can fall out of sync within a few commits. A runbook follows the same path, and the mismatch often becomes visible during an incident, when the steps on the page no longer describe the service in front of you.

Sumedh Masulkar, Senior Technical Lead at InMobi, agreed with the diagnosis and explained why the problem has resisted years of good intentions.

The Staleness Problem Has Outlasted Decades of Experience


The challenge is familiar even on experienced teams. Engineers know which documents matter, and they know what can happen when those documents fall behind. The difficulty is that updating the runbook usually happens after the code has already changed.

A pull request gets reviewed and merged. The service changes. The runbook waits for someone to revisit it later. By the time it gets updated, the system may already have moved on.

That is why reminders and documentation sprints only go so far. They improve the runbook for a point in time, but they don't change the workflow that causes it to drift.

Runbooks Became First-Class Citizens Alongside the Code

InMobi addressed the problem by changing where runbooks live. Runbooks, operational context, and event data now sit in the same GitHub repository as the service they describe.

That placement gives the documentation the same basic controls as the code: ownership, version history, pull requests, review, and a visible record of what changed. An engineer looking at the repository can see the implementation and the operating context together. A reviewer can assess both in the same change set.

The repository also gives the runbook a clearer relationship to the service. When a team owns the code, it owns the context checked in with that code. There is less ambiguity about which document is current, who should review it, or which version belongs to a given release.


Every Pull Request Checks Whether the Runbook Should Change

The next change happens inside the pull request, where the system already has the strongest signal that something material has changed.


The proposed update goes through the same review as the code itself. Engineers decide whether the change is accurate before it is merged, keeping documentation and code aligned through the same workflow.

By the time someone reaches for the runbook during an incident, it has already passed through the same review process as the service it describes.


The Pull Request Becomes the Documentation Checkpoint

Once documentation moves into the repository and the PR workflow, keeping it current stops depending on a separate maintenance habit. The update is prompted by the code change that created the need for it.

That creates a useful boundary for the agent. It does not need to roam through a wiki trying to infer which pages may be stale. It can focus on a known service, a known diff, and the runbook attached to that service. Reviewers can then judge the suggested update with the code change still fresh in context.

The same structure helps beyond the runbook itself. Operational events, service context, alert behavior, and troubleshooting steps can be versioned together. An AI system reading that material later receives context tied to the service and its history, instead of a loose collection of documents with uncertain ownership and freshness.

The End of Documentation Drift

Sumedh's approach starts with a simple observation: documentation drifts when its update path is separate from the work that changes the system. InMobi moved that path into the repository and made every pull request a checkpoint for the runbook and its surrounding context.

You can apply the same test to your own workflow. Pick a recent service change and trace what happened to the runbook. Was the documentation reviewed with the code? Did the change to metrics, alerts, dependencies, or operating steps appear in the same PR? Could a reviewer see the service change and the context change together? The answers show where drift is entering the process.

At StackGen, Aiden works with runbooks and operational context inside the engineering workflows your team already uses. Try Aiden for SRE Community Edition for free today.