Running the Agent Substrate demo on GKE
Copy as MarkdownIn the Agent Substrate post, I included a video of two interleaving agents running on Substrate. Here’s how I put it together.
The complete code is in WilliamDenniss/substrate-adk-demo. You’ll create a disposable GKE cluster, install Substrate with two timeout changes for long-running agent requests, and deploy the two ADK agents.
Check out both repositories
Start with the Substrate and demo repositories next to each other. The demo scripts look for Substrate at ../substrate by default:
mkdir substrate-adk-on-gke
cd substrate-adk-on-gke
git clone https://github.com/agent-substrate/substrate.git
git clone https://github.com/WilliamDenniss/substrate-adk-demo.git
Substrate moves quickly, so use this same checkout for the setup tool, installation, and kubectl-ate CLI. I initially had newer demo manifests talking to an older installed control plane; errors about spec.ateomImage or missing actor-template references are a sign of that version skew. For a disposable demo, a clean install from the current checkout is much simpler than trying to migrate the old resources.
Warning: Substrate is moving quickly with many breaking changes. In fact, in the time it took me to create this demo, there was a change that broke this demo. The current code is based on Substrate commit f936206dd544d2dcfd190fac61d1fd2fa0a933c0. If you check out Substrate HEAD, you’ll likely need to point your coding agent at this demo to refactor it, but the core principles should remain valid.
Install Substrate on GKE
Follow the GKE quickstart to install Substrate on a GKE cluster. Configure the development environment, then let the setup tool create the GKE cluster, GCS bucket and IAM bindings:
cd substrate
cp hack/ate-dev-env.sh.example .ate-dev-env.sh
# Edit .ate-dev-env.sh for your project, cluster, bucket and KO_DOCKER_REPO.
source .ate-dev-env.sh
gcloud auth application-default login --project="$PROJECT_ID"
go run ./tools/setup-gcp bootstrap
The setup tool creates the cluster with the Kubernetes beta certificate APIs Substrate requires. Those APIs must be enabled when the GKE cluster is created, so I would use the supplied bootstrap path for a first test rather than trying to retrofit an existing cluster.
KO_DOCKER_REPO must point to a container repository you can push to. The setup tool grants the required image-pull permissions, but it doesn’t create the repository itself. This is where ko publishes Substrate’s control-plane and gVisor worker images so GKE can pull them.
If you’re using a new Artifact Registry repository, create it once and configure Docker authentication before installing Substrate. This example matches the KO_DOCKER_REPO value used later in the demo:
gcloud artifacts repositories create ate-images \
--repository-format=docker \
--location="$GCE_REGION" \
--project="$PROJECT_ID"
gcloud auth configure-docker "$GCE_REGION-docker.pkg.dev"
Two timeout changes for LLM requests
I made two behavioral changes to manifests/ate-install/atenet-router.yaml to work better with this demo (allowing for longer LLM response times and more queuing). I recommend applying them before installing Substrate:
- Increase –parked-request-budget from its five-second default to one minute. An LLM turn can occupy all five workers for much longer than five seconds, so queued actors need more time for another actor to finish and suspend.
- Enable a five-minute –route-timeout. The normal end-to-end route timeout is too short for a model response that holds the HTTP request open throughout generation.
The relevant diff is:
spec:
template:
spec:
containers:
- name: atenet-router
args:
+ - "--parked-request-budget=1m"
# Long-running model calls hold the route open until generation ends.
- # - "--route-timeout=5m"
+ - "--route-timeout=5m"
These are demo values. A production value should come from your observed model latency and the amount of queuing you’re willing to accept. A longer queue can absorb a short capacity shortage, but it also means clients wait longer before they receive an error.
With those changes in place, install the system from the same checkout:
./hack/install-ate.sh --deploy-ate-system
How the demo is wired
Both containers run ADK’s standard API server. The ActorTemplates override the images’ command with the equivalent of:
CMD ["adk", "api_server", "--host=0.0.0.0", "--port=80", "/app"]
Port 80 is Substrate’s default HTTP ingress port for actors. The absolute /app is also deliberate: the current runtime starts the process with / as its working directory, so using . makes ADK search the wrong directory and the eventual /run call returns 404.
The news scout agent container image exposes the hn_signal_scout ADK app, while the earthquake agent container image exposes quake_agent. The two images are already published as docker.io/wdenniss/hnscout:latest and docker.io/wdenniss/quakeagent:latest; the Artifact Registry repository configured during setup is for Substrate’s images, not these agent images. (Hopefully once Substrate is a little more mature there will be released artifacts you can use directly).
Both templates are configured to use Full snapshots, which capture process memory, the changed root filesystem, and durable-volume data. That’s important for ADK because its in-process session service and local artifacts survive the suspend/resume cycle. Substrate also supports Data snapshots for applications that only need durable volume contents and can cold boot the process again. The GCS bucket configured during setup stores the golden and per-actor snapshots.
Image: Golden and per-actor snapshots stored in the configured GCS bucket
Both agents call Gemini through Gemini Enterprise Agent Platform (formerly known as Vertex AI). Their Google client libraries use Application Default Credentials (ADC), which on GKE resolves to short-lived credentials supplied through Workload Identity Federation for GKE. There is no GOOGLE_API_KEY in the templates. Instead, the workload proves its identity and receives short-lived access tokens rather than carrying a reusable bearer secret. On Google Cloud, IAM and Workload Identity can be configured as part of deployment, so the application obtains credentials automatically. The demo grants the shared Substrate egress identity permission to call Gemini Enterprise Agent Platform.
Matching labels let every worker run an actor from either template. Substrate automatically resumes an actor when traffic arrives, but it doesn’t decide when an arbitrary HTTP application is idle. Completing an ADK response therefore doesn’t automatically suspend the actor. The demo runner explicitly suspends each actor after every model response; a real agent harness needs to do the same, or implement its own idle policy. The counter demo also suspends explicitly—this isn’t an automatic side effect of finishing an HTTP request.
Configure and run the agents
Move into the demo checkout and create its local environment file:
cd ../substrate-adk-demo
cp .env.example .env
Set these values in .env, using the same project, bucket and image repository you configured for Substrate:
GOOGLE_CLOUD_PROJECT=your-project-id
GOOGLE_CLOUD_LOCATION=global
BUCKET_NAME=your-substrate-snapshot-bucket
KO_DOCKER_REPO=us-west1-docker.pkg.dev/your-project-id/ate-images
Now grant the demo’s model permission and validate the configuration:
make grant-auth
make validate
The make targets are just short names for shell scripts; Substrate doesn’t depend on Make. I find make actor-demo easier to remember, but ./scripts/actor-demo.sh does the same thing.
Keep the router private and port-forward it in another terminal:
kubectl port-forward -n ate-system service/atenet-router 8000:80
With that terminal running, the concrete demo lifecycle is four commands. Run them one at a time; after make deploy is a good time to open the monitoring terminals in the next section.
make deploy
make actor-demo ACTORS_PER_AGENT=7
make actor-cleanup
make cleanup
make deploy creates the five-worker pool, creates both ActorTemplates, and waits for both golden snapshots. It doesn’t create any ordinary actors, so all five workers are idle when deployment finishes. The wait for a golden snapshot is real work—the agent container has to boot, pass its health check and be checkpointed—so it can take a little while. If the template enters a failed state, inspect that error rather than waiting indefinitely.
The second command creates seven Hacker News actors and seven earthquake actors, then submits all fourteen two-turn workflows concurrently. The five-worker pool runs five actors at a time while the router parks the excess requests. Each actor receives a different first prompt and a follow-up that depends on the first result.
make actor-cleanup removes those fourteen numbered actors but keeps the templates and worker pool. Finally, make cleanup removes the remaining demo resources. This gives you a useful pause between the third and fourth commands to confirm that selective actor cleanup really did leave the infrastructure alone.
If you omit ACTORS_PER_AGENT, the starting value is five per agent (ten actors total). You can also limit how many workflows the client submits at once:
make actor-demo ACTORS_PER_AGENT=7 MAX_CONCURRENCY=8
MAX_CONCURRENCY doesn’t change the WorkerPool. If client concurrency is no higher than worker capacity, there won’t be excess requests for the router to park.
If you want a smaller check after make deploy, run make smoke. It creates or reuses one actor from each template, sends one real model prompt to each, prints the results, then suspends both actors.
What to watch
This demo is more interesting with a few terminals open before you run make actor-demo. First, install kubectl-ate as a kubectl plugin if you haven’t already:
cd ../substrate
go install ./cmd/kubectl-ate
cd ../substrate-adk-demo
If kubectl says unknown command “ate”, the resulting kubectl-ate binary isn’t on your PATH. You can call kubectl-ate directly, or add your Go binary directory to PATH.
Now watch the physical Pods, logical actors, and workers:
watch -d 'kubectl get pods -n ate-demo-python-adk'
watch -d 'kubectl ate get actors -a ate-demo-python-adk'
watch -d 'kubectl ate get workers'
On macOS, watch is available from Homebrew (brew install watch). As the demo runs, only five workers can be assigned at once. Actor states move between RUNNING and SUSPENDED, and a later resume may place an actor on a different worker. Once the demo finishes, every numbered actor should be suspended and all five workers should be free.
To see request parking, expose the router’s status port in another terminal:
kubectl -n ate-system port-forward deployment/atenet-router 4040:4040
Then watch the active parked-request count while make actor-demo is running:
watch -d "curl -fsS 'http://127.0.0.1:4040/statusz?format=json' | jq -c '.parking'"
The useful field is .parking.active. It should rise while all workers are occupied and fall as actors finish, suspend, and release capacity. If you look only after the run, it will quite correctly be zero.
The demo runner prints the actual model responses. In my run, one earthquake actor found the deepest event, suspended, then handled this follow-up after it was restored:
Prompt: Convert the depth you just reported from kilometers to miles and show the calculation.
Result: 64.644 km × 0.621371 = 40.167 miles
The exact earthquake and Hacker News results will of course change. The useful part is that the follow-up understands the depth you just reported even though the process was snapshotted and the worker released between turns.
For the container’s stdout and stderr while an actor is active, use the actor-aware log command:
kubectl ate logs actors quakeagent-4 -a ate-demo-python-adk -f
That follows the actor rather than a particular Pod. A suspended actor has no active Pod to query, so for this short demo the labeled API responses printed by make actor-demo are usually the easiest view. Substrate also adds actor, atespace, template, and container metadata to structured logs for a centralized logging backend; the observability guide covers that model in more detail.
You can also scale the warm worker capacity using the normal Kubernetes scale subresource:
kubectl scale workerpool/python-adk \
-n ate-demo-python-adk \
--replicas=8
kubectl rollout status deployment/python-adk -n ate-demo-python-adk
Run the demo again and you’ll see less parking because eight actors can be active at once. A later make deploy reapplies the demo manifest’s default of five workers.
Cleanup
To remove only the numbered actors created by make actor-demo, while keeping the templates, pool, and any suspended smoke-test actors:
make actor-cleanup
To remove the full demo and revoke its Gemini Enterprise Agent Platform permission:
make cleanup
make revoke-auth
The cleanup deliberately retains the snapshot objects in GCS, so delete that prefix or tear down the disposable project when you’re done. The shared atenet-egress identity used here is also intentionally demo-only: every actor routed through it can obtain the same Google Cloud credentials. A production design needs actor-aware credential isolation, and the unauthenticated ADK API server should never be exposed publicly as configured here.