GLOBAL ENTERPRISE

How a global enterprise is Building the Autonomous SRE Future with Aiden

Customer extended NOC-only monitoring with StackGen's OTel-based managed observability ObserveNow, extending Observability access from ~5–10 NOC engineers to ~200 developers across the org, and layered in Aiden as an AI SRE copilot.

The Problem: Scaling SRE Without
Scaling Headcount

As this enterprise began the journey to modernize their internal platforms, these were the challenges they faced:

Observability was restricted to only NOC teams. Tools like Zabbix and Site24x7 were accessible only to the NOC team. Developer scrum teams had less visibility into their own deployments — they had to manually request log snippets from the NOC to debug issues.

The signal-to-noise ratio was creating toil. The NOC team was fielding ~140 alerts per day, with only 20–30 being actionable. Human triage was consuming valuable engineering hours.

Global dev teams required global ops support. With engineers across the US, Europe, India, and Ukraine, 24/7 cloud ops coverage was a challenge. A UAT promotion that should take seconds could block a Scrum team for hours if no one with access was online.

Compliance constraints limited self-service. SOC2, PCI, HIPAA, and GDPR requirements meant Scrum team developers could not be given deployment permissions - creating a bottleneck that required cloud ops intervention for every promotion to UAT and production.

Heterogeneous version control. Code lived across Bitbucket, Azure DevOps, GitHub.

"The Scrum teams had no access at all. If there was a critical incident, the NOC had to take a cut-and-paste of the server log and send it to them. Observability should mean exactly that — observable — and it should be democratized."

VP of Cloud Infrastructure and Operations, Global Enterprise

"We are not in the business of building an observability stack or SRE stack. Let's focus on our business and use StackGen [Aiden]."

VP of Cloud Infrastructure and Operations, Global Enterprise

The Solution: A Layered, Phased Deployment

The enterprise rollout moved through two deployment phases toward a longer-term autonomous vision.

Phase 1 — Democratizing Observability with Aiden for Observability

The enterprise deployed StackGen's managed observability ObserveNow as a private SaaS instance — with OpenTelemetry (OTel) collectors running in every Kubernetes (EKS) cluster across dev, test, UAT, and production environments.

The choice of OTel as the instrumentation standard was deliberate:

Common grammar: OTel provides a language- and vendor-agnostic standard for traces, logs, and metrics, making adoption consistent across Java, C#, and other stacks.

No vendor lock-in: Engineers instrument their code once and remain free to route signals to any backend.

Democratization by design: Unlike Zabbix (NOC-only) or Site24x7 (license-capped), ObserveNow offers tiered access — every engineer gets a viewer role, so the full observability picture is available to whoever needs it.

Adoption approach: the Ambassador Program
Rather than a top-down rollout, the enterprise built a grassroots adoption model:

POC with engineering leaders. Architects, CTOs, and principals saw a live demonstration of ObserveNow's APM, distributed tracing (waterfall views), and log search capabilities.

Ambassador seeding. A small group of principal engineers became internal champions, evangelizing ObserveNow within their Scrum teams.

Office hours and brown bags. Regular sessions run jointly by StackGen and the enterprise’s cloud ops team gave engineers hands-on experience.

Video training library. Short, screen-recorded tutorials from StackGen team covering OTel basics, ObserveNow components, log searches, and APM dashboards served as an always-available onboarding resource.

AI-generated user guide. The enterprise used Aiden itself to generate a contextualized ObserveNow user guide — a compelling early demonstration of Aiden's value to skeptical engineers.

Outcome: ~200 engineers across the enterprise’s global engineering organization now have access to ObserveNow. Power users actively attach APM trace links and waterfall evidence to bug reports, significantly reducing back-and-forth in incident tickets.

Phase 2 — AI-Assisted RCA with Aiden

With the observability foundation in place, global enterprise began integrating Aiden — StackGen's AI SRE agent — layered on top of ObserveNow, Argo CD, and GitHub.

Current use case: Engineers experiencing test failures or deployment issues query Aiden directly. Aiden investigates across the connected infrastructure, checking Kafka permissions, S3 bucket configurations, pod states, and more, and provides a root cause summary with a shareable link that engineers attach to their Jira tickets.

Global enterprise ran a controlled autonomous Aiden POC on dev environments, targeting a high-frequency, well-understood failure mode: out-of-memory (OOM) crashes and pod restart loops.

1. Aiden was trained to detect OOM signals and validated its predictions with a built-in eval loop, a double-check before taking action.
2. When predictions were confirmed, Aiden autonomously increased memory allocation by 10% and restarted the affected pod.
3. Accuracy: 90–95% on predicting and resolving OOM issues in the dev environment.

This controlled experiment established the trust-building model the enterprise will apply to all future autonomous capabilities: POC in dev, validate accuracy, earn stakeholder trust, expand scope.

"If a test fails, I ask Aiden: 'Why did that invoice update test fail?' Aiden says, 'It failed because your Kafka permission got changed.' They add that Aiden investigation link to the ticket. That's the new workflow."

VP of Cloud Infrastructure and Operations, Global Enterprise

"Aiden is global and never sleeps. If I have a never-sleeping global team, let them talk to a never-sleeping agent."

VP of Cloud Infrastructure and Operations, Global Enterprise

Safe Autonomous promotion

Today, when a Scrum team in APAC finishes a tested microservice at midnight Pacific time, they cannot promote it to UAT — compliance rules prohibit developers from self-promoting code. Someone from cloud ops must manually click a button. If no one is available, delivery is blocked.

Customer’s vision: Scrum teams ask Aiden, Aiden checks that all preconditions are green (QA approval, compliance guardrails, CI/CD pipeline status), and — with appropriate service-user permissions — Aiden executes the promotion. No human bottleneck. No VP manually clicking a GitHub button at 1 AM.

Aiden's role-based governance model maps directly to customer's SOC2, PCI, HIPAA, and GDPR requirements. Developers write feature code; Aiden (like cloud ops) does not. Aiden may generate Terraform for infrastructure, but never touches application feature code. Separation of duties is preserved, and auditors get the controls they need.

Results

Metric Before StackGen With StackGen
Engineers with observability access ~5–10 (NOC only) ~200 (all engineering)
Observability tools Zabbix + Site24x7 (siloed, license-capped) ObserveNow (OTel-native, role-based access)
Incident debugging workflow Manual NOC log handoff to Scrum teams Self-serve APM + Aiden RCA with shareable evidence links
Autonomous RCA accuracy (OOM, dev) N/A 90–95%
CI/CD standardization Fragmented (Jenkins, Bitbucket, SVN, Perforce) Unified GitHub Actions + Argo CD (managed by StackGen)
SLA target 99.95% uptime (22-min monthly error budget)
All

Start typing to search...