Agents do not behave like web services. They sit idle waiting on a model, then burn money in a tight loop. They accumulate state, call outbound APIs, and need a sandbox you can pause without losing the thread. Kubernetes was not designed for that shape of work. Google’s answer is AX, an open-source Google AX agent orchestrator that treats agents as suspendable actors instead of pods that must stay warm.
v0.3.0 shipped on 20 September 2026. It hit the top of Hacker News the next day. InfoQ, AI Beat, and daily.dev all covered the same four primitives: Task, Workspace, Gateway, Model. What they did not ship is a walkthrough you can actually type. That is this post.
I spent this morning on the official repo at github.com/google/ax, the concepts doc, the runner contract, the sandbox metadata server, and the exact YAML in examples/. I did the work a developer does on day one.
Official source: github.com/google/ax
Why agents need their own orchestrator, not just Kubernetes

A microservice request starts, finishes, and dies. An agent plans, waits on an LLM, retries a tool, asks a human, then resumes an hour later. If you keep that process alive as a normal container, you pay CPU and memory while it sleeps. If you kill it, you lose the working memory the next turn needs.
AX states the problem in the README without marketing fog. Agents accumulate state. They need isolation because the code they generate is untrusted. They call model APIs and tool servers. They can burn money in a loop if nobody is watching.
The control plane answers with four small resources under ax.io/v1alpha1. You declare them in YAML. You apply them with a CLI that copies kubectl verbs on purpose: apply, get, describe, watch, delete, plus suspend, resume, and ssh.
v0.3.0 also moved task state off Kubernetes custom resources and onto Redis Streams. That is the quiet architectural tell. etcd was never built for millions of short-lived agent tasks spinning up and dying. Redis Streams is.
The four primitives of the Google AX agent orchestrator
A Task is the smallest unit of isolated execution. It names an image, a command, CPU and memory limits, a Gateway, and one or more Workspaces. The first workspace becomes the working directory. AX does not try to model the whole agent graph. It gives you one cheap unit you can create, isolate, freeze, and throw away.
A Workspace is the warm start. Clone Git repos, attach MCP servers, materialize skill packages, and optionally hand a plain-language goal to a bootstrap agent so the toolchain exists before the real command starts. Declare the Workspace once. Bind it from many Tasks.
A Gateway is the network fence. You write an explicit host allowlist. An untrusted coding agent should not have host: "*" on 443 in production, even if the example YAML does.
A Model is the LLM config for the platform itself, with credentials pulled from a Kubernetes secret. Rotate a key once. Every Task that references the Model picks up the new credential on the next apply.
Suspend and resume sit next to those four. Idle agents checkpoint and leave the CPU. Resume is designed for sub-second wake-up with no cold start tax. ax ssh lets you look over the agent’s shoulder when spec.debug is true.
Hands-on: install and run your first AX task in 10 minutes
This is the 10-minute path from the official README. You need Go on your PATH, a Kubernetes cluster, ko (brew install ko), a container registry the cluster can pull from, Redis (the deploy target puts it in ax-system), and a reachable Agent Substrate control API. The in-cluster default is api.ate-system.svc.cluster.local:443.
Install the CLI:
go install github.com/google/ax/cmd/ax@latestThat drops ax in $(go env GOPATH)/bin. Confirm that directory is on PATH before you blame the binary.
Deploy the control plane:
make deploy AX_IMAGE_REPO=<your-registry>This builds images with ko and lands Redis plus the three v0.3.0 services (API front end, reconciler, sandboxed task runner) in the ax-system namespace.
A minimal multi-document file looks like the README example:
apiVersion: ax.io/v1alpha1
kind: Workspace
metadata:
name: golang
spec:
git:
- repo: https://github.com/golang/go.git
branch: "my-fix"
---
apiVersion: ax.io/v1alpha1
kind: Task
metadata:
name: test
spec:
workspaces:
- name: golang
goal: "Ensure that Go tool chain is available and is built from source"
debug: trueThen the loop you will live in:
ax apply -f examples/task.yaml
ax get tasks
ax watch task test
ax ssh test -- ls -al /workspace
ax suspend task test
ax resume task testdebug: true is required for ax ssh. Without it, guest services stay off because they allow arbitrary process execution inside the sandbox. That is the correct default.
The runner is always /usr/local/bin/ax-task-runner, even if you set a custom image. The controller never uses spec.command as the container entrypoint. Your image must contain that binary or a wrapper at that path. The runner reads AX_TASK_YAML and AX_WORKSPACES_YAML, prepares each workspace, serves /healthz and /readyz on port 80, then starts your command as a child with the first workspace as cwd.
If a Workspace binding has a goal, a bootstrap agent runs first and the Task stays not-ready until that finishes. Default timeout is 10 minutes. Set AX_BOOTSTRAP_TIMEOUT if your toolchain build is slower. The bootstrap agent expects GEMINI_API_KEY in the container unless you skip goals and bake the image yourself.
Want the full lifecycle without inventing YAML? The repo ships ./demo.sh. It applies a custom workspace, waits for readiness, runs commands over ax ssh, and suspends the task.
Google AX vs Kubernetes for agents
Kubernetes still wins at cluster plumbing. Nodes, networking, secrets, RBAC, and the ecosystem around kubectl are why AX speaks that language instead of inventing a new one. ax follows your active kube context. Switch with kubectx and AX tunnels to that cluster’s control plane.
Where Kubernetes alone is the wrong tool is the cost model. A Deployment keeps replicas warm. An agent fleet spends most of its life waiting. AX’s bet is dense multiplexing: checkpoint idle actors, pack many Tasks onto shared workers, resume in under a second. Agent Substrate claims 10x higher density and sub-500ms resume at high activation rates — treat those as vendor claims until you measure them.
Use Kubernetes when the unit of work is a long-lived service. Use the Google AX agent orchestrator when the unit of work is a stateful, untrusted, pauseable agent that should not hold a full pod tax while it waits on tokens.
Claude Code, Cursor, and ZCode still win for a single developer on a laptop. AX starts to make sense when you need many agents, hard isolation, an egress allowlist, and a platform team that wants YAML instead of a proprietary Agents API.
Isolation, suspend, and the cost reality check
Three things I marked in the docs that most launch posts skip.
First, early stage is not a footnote. The README warns that core concepts and protocols will take breaking changes before a stable release. External PRs are paused while the team stabilizes the architecture. Build against AX today if you want to learn the model. Do not bet a production SLA on v0.3.0 APIs staying still.
Second, the example Gateway is wide open. host: "*" on 443 is a tutorial convenience. Production Gateways should name the model host, the Git host, and the MCP hosts you actually need. Then test a blocked destination on purpose so you know the fence works.
Third, AX is free under Apache-2.0. The bill is operational: Kubernetes, Redis, a registry, LLM API keys, and the tokens the agents spend. Suspend/resume only helps if your harness actually goes idle.
Self-hosting is possible off GKE. Portable is not plug-and-play. You still need Substrate’s control API, the runner binary at the fixed path, and a cluster that can pull the images ko produces.
When to reach for AX, and when to stay with Claude Code
Stay in Claude Code when you are one engineer, one repo, and the agent lives in your terminal. CLAUDE.md, AGENTS.md fallback, skills, and hooks already solve that job.
Reach for AX when you outgrow one machine. You want a platform primitive for agent Tasks. You need to ssh into a runaway sandbox at step 400. You need to freeze 200 idle agents overnight without deleting their workspace. You want MCP servers and Git checkouts declared once, not copied into every harness container.
A useful hybrid: keep Claude Code as the authoring surface. Use AX as the runtime when a Task must run unattended in a cluster with a Gateway in front of it.
Production gotchas I hit on day one of reading the contract
The working directory used to leak. Older runners started tool calls from / instead of the workspace. v0.3.0 made the workspace directory authoritative and added WorkspaceSystemInstruction so the agent emits workspace-relative paths. If a tool still writes to /tmp or /, your custom runner is ignoring WorkDir.
ax delete blocks until teardown finishes. Do not script a fire-and-forget delete and assume the sandbox is gone.
ax talks to the control plane over gRPC, not the Kubernetes API, for Task lifecycle. ax ctx shows which kube context it is using and how it reaches the control plane. When ax get tasks is empty and kubectl get pods -n ax-system looks healthy, you are pointed at the wrong namespace or the wrong context. Default namespace is default. Flag is -a.
Goal-based Workspace bootstrap needs a model key in the sandbox. If you do not want a second agent touching your cluster on first boot, omit goal and bake the toolchain into the image.
Common questions about Google AX
What is Google AX?
AX is Google’s open agentic orchestrator. You declare Tasks, Workspaces, Gateways, and Models as Kubernetes-style YAML. The runtime sandboxes the work, pre-wires Git and MCP, fences egress, and can suspend an idle agent then resume it without a cold start. The project lives at github.com/google/ax under Apache-2.0.
How does AX differ from Kubernetes?
Kubernetes schedules containers. AX schedules agent work on top of Kubernetes and Agent Substrate. The extra verbs are the point: suspend, resume, ssh into a live sandbox, and bind a Workspace so every Task starts warm.
Can I self-host Google AX without GKE?
Yes in principle. The docs describe any Kubernetes cluster plus Redis, a registry, ko, and a reachable Agent Substrate control API. Expect integration work, not a one-line SaaS signup.
Does AX work with Claude or MCP?
Workspaces can declare MCP servers and skill registries. Claude Code remains a separate harness. You can run a Claude-oriented command inside a Task if your image and Gateway allow the Anthropic API host.
How do I suspend and resume an agent task?
Use these commands:
ax suspend task <name>
ax resume task <name>Use this when the agent is waiting, not when it is mid-write to a file you cannot snapshot cleanly. Measure resume latency on your own hardware before you quote the sub-second claim in a design doc.
1 comment