Blog

KubeCon NA 2026: The Production Stories Worth Hearing

Written by Neel Shah | Oct 8, 2026, 10:12:33 AM

KubeCon + CloudNativeCon North America 2026 lands in Salt Lake City on November 9–12. Here's how we're planning our week, and the sessions we've already circled.

In five weeks, the cloud native community heads back to Salt Lake City. KubeCon + CloudNativeCon North America 2026 runs November 9–12 at the Salt Palace Convention Center, and the schedule is finally live.

We're excited for a simple reason. This year's program is full of production stories. There are talks about a GitOps commit that deleted 39 clusters, a time skew that locked a cluster for a day, and a CNI migration across hundreds of clusters with zero downtime. Those are the problems platform engineers and SREs deal with every week.

AI is also everywhere on the agenda, from GPU scheduling to agent observability. What we like is that the AI sessions are grounded in the same fundamentals as the rest of the program: how a workload runs, scales, fails and recovers.

Below is our guide to the talks worth your time, organised by the problem you're trying to solve.

KubeCon NA 2026 at a glance 

The main conference runs Tuesday to Thursday, with Monday given over to co-located events and Project Lightning Talks. There is no virtual component this year, so the hallway track only happens in person. Keynote and breakout recordings land on the CNCF YouTube channel within about two weeks.

Day

What's on

Mon, Nov 9

CNCF-hosted co-located events and Project Lightning Talks

Tue, Nov 10

Keynotes, breakouts, Maintainer Track, ContribFests, Solutions Showcase; #KubeCrawl + #CloudNativeFest 6:00–7:30 PM

Wed, Nov 11

Keynotes, breakouts, Maintainer Track, ContribFests, Solutions Showcase

Thu, Nov 12

Keynotes, breakouts, Maintainer Track, ContribFests for Argo CD, Meshery and others

 

Which pass? The All-Access Pass covers Monday's co-located events plus the three main days. The KubeCon + CloudNativeCon Only Pass covers the main conference and Monday's Lightning Talks, but not the co-located events. If any of the Monday picks below matter to you, go All-Access. Current pricing is on the registration page.

Seating is first come, first served, so build your agenda in the online schedule before you fly.

Which Monday co-located events are worth it? 

Monday is where you go deep on one community. This year's lineup includes Agentics Day: MCP + Agents, ArgoCon, BackstageCon, CiliumCon, Cloud Native AI + Inference Day, FluxCon, Kubernetes on Edge Day, Observability Day, Open Source SecurityCon, OpenTofu Day and Platform Engineering Day. You can only be in one room at a time, so pick by the problem on your desk. 

If your week is mostly…

Start Monday at

Incidents, alerts and telemetry

Observability Day — OpenTelemetry, Prometheus, Fluent Bit, Jaeger, Thanos, Perses and more

Building an internal developer platform

Platform Engineering Day

Terraform and infrastructure as code

OpenTofu Day

GitOps and continuous delivery

FluxCon or ArgoCon, where the community is looking toward Argo CD 4.0

Networking and eBPF

CiliumCon — Cilium, Hubble and Tetragon

Agents and AI workloads

Agentics Day: MCP + Agents, or Cloud Native AI + Inference Day


 Our own team will be split between ArgoCon, Observability Day and Agentics Day. AI agents that touch production need the same telemetry and guardrails as everything else, and those two rooms are where that conversation is happening. 

Don't miss: StackGen at ArgoCon 

Scheduling AI Workloads on Kubernetes with Argo: Patterns from Production Neel Shah, Developer Advocate, StackGen · Monday, Nov 9, 5:00–5:10 PM · Salt Palace, Level 2, Room 251 A-C · ArgoCon Lightning Talk, any level

AI pipelines are as much a scheduling problem as a modelling one. Training jobs compete for expensive GPUs, batch inference has to coexist with real-time serving, and one failed step can burn hours of compute. Neel will share production patterns for running AI pipelines with Argo Workflows: GPU-aware DAGs, recovery for long-running training jobs, model artifact versioning across pipeline steps, and human approval gates before a pipeline executes.

It's ten minutes, so arrive early. ArgoCon is a co-located event, so you'll need an All-Access Pass.

Must-see talks for SREs: reliability and incident stories 

The best KubeCon talks are honest postmortems. These are the ones we expect to fill up first.

GitOps Meltdown: How a Commit Triggered a Cascading Delete of 39 Clusters in Production One commit, 39 clusters gone. Anyone running GitOps at scale should hear how the blast radius got that big and which safeguards would have stopped it. Expect lessons on pruning, sync policies and change review.

When the Clock Lies: How an 8-Minute Time Skew Locked Kubernetes Cluster for a Day Time skew is the kind of root cause that hides behind a dozen misleading symptoms. This is a good reminder of why incident investigation has to follow the evidence rather than the first plausible alert.

Green Dashboards, Broken Clusters: Taming Stale Watches in Kubernetes Operators Every on-call engineer knows the feeling of a dashboard that says everything is fine while users say otherwise. This session gets into why operators drift from reality and how to detect it.

Hidden Pod Startup Latency: Container Image Pull Slow pod startup quietly erodes autoscaling and recovery time. A practical look at one of the most overlooked contributors to MTTR.

Battle-tested Autoscaling Paradigms with KEDA Event-driven autoscaling patterns that have survived real traffic, from the KEDA community.

Filter, Transform, Route: Mastering OTTL in the OpenTelemetry Collector If telemetry volume or cost is on your list, OTTL is how you take control of it in the Collector. Bring your noisiest pipeline in mind.

Must-see talks for platform engineers: infrastructure at scale 

Platform teams are being asked to run more clusters, more workload types and more automation without losing reliability. These sessions tackle that directly.

Zero Downtime CNI Migration at Scale: Canal to Cilium Across Hundreds of Production Clusters Our pick for the talk of the week. A network-layer migration across hundreds of live clusters is the kind of change most teams postpone for years. Come for the rollout plan and the rollback story.

When Infrastructure Breaks in Production: How EarnIn Brought Testing to the Full CNCF Stack Most teams test application code and trust the platform underneath. EarnIn tests the platform too. Useful for anyone who has watched an infrastructure change take down a healthy service.

The New Scaling Matrix: Balancing In-Place Resizing, VPA, HPA and Automation Solutions In-place pod resizing changes the autoscaling conversation. This session helps you decide which lever to pull and when.

From Pending Pod to Ready Node: Declarative Worker Images for Kubernetes Scale-Out Node provisioning time is part of your recovery time. A look at treating worker images as declarative, versioned infrastructure.

RBAC as a Platform Capability: Applying Platform Engineering to Kubernetes Access Access control is usually a ticket queue. This talk frames it as a self-service platform feature with guardrails, which is where governance belongs.

Must-see talks on AI infrastructure and agents 

AI is the biggest new thread in this year's program. CNCF added a dedicated AI Inference + Agentic track, and Maintainer Track sessions cover AI-aware networking, GPU device management and AI/ML workload orchestration.

Follow the GPUs: How a Regulated Fortune 100 Serves Customers When No Single Cloud Has Enough GPU scarcity is now an architecture constraint. A regulated enterprise explains how it spreads AI workloads across clouds without losing control.

Many Clusters, One GPU Pool: Locality-Aware Scheduling and Data Delivery With Kueue and Dragonfly Treating GPUs across many clusters as a single pool, with data delivered close to where jobs land. Strong material for anyone planning a shared AI platform.

The question we'll be asking in every AI session is a reliability one. When an agent can act on live infrastructure, who approves the change, what does it know about the system, and how do you verify the result? Our own research found at least nine documented cases in the past year of an AI agent taking destructive action against live production on its own (StackGen State of Reliability 2026). That's why we think agent observability and policy deserve as much stage time as inference speed.

Beyond the sessions 

Recordings will be online in two weeks. What you can't watch later is the conversation after the talk, so leave white space in your agenda.

  • Maintainer Track. Hear roadmaps straight from the people building Kubernetes SIGs, OpenTelemetry, Prometheus, Cilium, etcd, Flux, Kyverno, OpenCost and more.
  • ContribFest. Hands-on sessions with maintainers for Harbor, OpenTelemetry, Podman, Argo CD, Meshery and others, several of them open to first-time contributors.
  • Project Pavilion. Bring the production problem your team has fought for six months and ask the maintainers directly.
  • #KubeCrawl + #CloudNativeFest. Tuesday 6:00–7:30 PM in the Solutions Showcase, and the easiest way to meet people outside your usual circle.

A tip from past KubeCons: arrive with one real problem in mind, such as a recurring incident, a migration or a scaling limit. Use it to filter the schedule, and you'll leave with answers rather than a bag of stickers.

Meet StackGen in Salt Lake City 

We'll be at KubeCon all week and would love to compare notes. Find us at Booth #283 in the Solutions Showcase, and catch Neel Shah's ArgoCon lightning talk on Monday at 5:00 PM in Room 251 A-C.

Our focus this year is the gap between AI agents and production reliability. Aiden for SRE handles alert triage, incident investigation, root cause analysis and policy-bounded remediation. It works with the telemetry stack you already run, and remediation is proposed as a pull request for your team to approve rather than applied silently. Aiden for InfraOps covers drift detection and IaC lifecycle across your existing tools.

Three things to take away before you build your agenda:

  1. Pick Monday's co-located event by the problem you're solving, then choose the All-Access Pass if you need it.
  2. Prioritise the postmortem-style talks. They fill up fastest and teach the most.
  3. Treat AI sessions as reliability sessions, and ask how agents are governed in production.

Want a head start? Book time with our team at KubeCon, or try Aiden for SRE Community Edition, free for up to two users. See you in Salt Lake City.

FAQ 

When and where is KubeCon North America 2026? November 9–12, 2026, at the Salt Palace Convention Center in Salt Lake City, Utah. Monday is co-located events; the main conference runs Tuesday to Thursday.

How many people attend KubeCon NA? Organisers and industry trackers expect roughly 9,000–12,000 attendees.

Is there a virtual option? No. Keynote and breakout recordings are posted to the CNCF YouTube channel within about two weeks of the event.

Does the KubeCon + CloudNativeCon Only Pass include co-located events? No. It covers the three main days and Monday's Project Lightning Talks. You need an All-Access Pass for Observability Day, Platform Engineering Day, Agentics Day and the other CNCF-hosted co-located events.

What are the best KubeCon 2026 talks for SREs? Start with GitOps Meltdown, When the Clock Lies, Green Dashboards Broken Clusters and the Canal-to-Cilium migration talk. All four are drawn from real production incidents.