Skip to main content
Version: v2.0

Prepare a Kubernetes Deployment for Production

Deploy on Kubernetes gets a working release into your cluster. This page covers what you change before other people depend on it.

The sections below are independent. Read the ones that describe your cluster.

Community-Maintained Beta

The Agenta Helm chart is community-maintained and currently in beta. If you encounter issues or have suggestions, please open a GitHub issue or reach out in our Slack community.

Survive a node going away​

Kubernetes moves pods off a node whenever the platform needs that node. A node upgrade, a node repair, and a scale-down that packs pods onto fewer nodes all do this. On GKE Autopilot it happens with no warning, and it is not a fault.

A workload with one replica is unreachable while its pod starts somewhere else. On one Agenta stage cluster that took about three minutes.

Step 1: give each public workload two replicas​

A second replica is the only thing that makes a workload survive the loss of its node. No other setting in this section does that.

api:
replicas: 2
web:
replicas: 2
services:
replicas: 2

Two replicas are enough for the chart to add a PodDisruptionBudget and topology spread constraints to that workload. Both are described in the next step.

replicas works on api, web, webMobile, services, workerStreams, workerQueues and supertokens.

Do not give the agent runner two replicas

agentRunner.replicas exists, but the API does not route a session to the runner pod that owns it. Two runner pods disagree about who owns a session. The chart defaults the runner to strategy: Recreate for the same reason.

Four workloads ignore replicas and always run one pod. cron runs supercronic, which fires every schedule in every pod. redisVolatile is a cache, and two caches answer differently. redisDurable and store.seaweedfs hold state in a single volume. Each of those four is unavailable while its pod restarts.

webMobile has no enabled switch. The mobile web app always deploys, and /m always routes.

Step 2: read back what the chart added​

The chart renders two objects per workload, and both are on by default.

A PodDisruptionBudget tells the platform how many pods of a workload may be down at once for a voluntary disruption. A voluntary disruption is one the platform chooses, such as a node drain. A node that fails outright is not voluntary, and no budget applies to it.

topologySpreadConstraints keep the replicas of one workload on different nodes and in different zones.

ReplicasPodDisruptionBudgettopologySpreadConstraints
1none, unless you set <workload>.pdb.protectSingleton: truenone
2 or moremaxUnavailable: 1one pod per node (hard), even spread over zones (soft)

The generated constraints are scoped to one Deployment revision. A rolling update is therefore not blocked by the pods it replaces.

Check both after an upgrade:

kubectl -n agenta get poddisruptionbudgets
kubectl -n agenta get deployment agenta-api \
-o jsonpath='{.spec.template.spec.topologySpreadConstraints}'

What the defaults do not give you​

Four limits apply:

  • A budget over a single pod does not keep that workload up. It defers the eviction of its node, then the pod still moves.
  • GKE Autopilot respects a budget for one hour during a node upgrade and then proceeds anyway.
  • No budget applies to a node that fails outright.
  • The bundled PostgreSQL gets a maxUnavailable: 1 budget from the Bitnami subchart, over a single primary. That is the deferral case above. Use a managed database if the deployment must survive the loss of the database node. See External PostgreSQL.

Step 3: decide what to do about the single-pod data stores​

redisVolatile, redisDurable and supertokens carry cluster-autoscaler.kubernetes.io/safe-to-evict: "false". A cluster autoscaler then leaves their node alone during consolidation. The annotation does not stop an operator drain.

On GKE Autopilot the annotation also delays automatic node upgrades for that node, and GKE can evict the pod anyway after about seven days. The annotation is therefore a delay, not a guarantee. Back up the single-pod data stores whatever the annotation says.

To let the autoscaler consolidate the node of one of them:

redisVolatile:
safeToEvict: true

To defer the eviction of a single pod that has no replica, ask for a budget over it:

agentRunner:
pdb:
protectSingleton: true

Step 4: override a budget or a constraint, if you need to​

<workload>.pdb.maxUnavailable and <workload>.pdb.minAvailable replace the generated budget. Set one of the two, not both. <workload>.topologySpreadConstraints replaces the generated constraint list with the list you give.

api:
replicas: 3
pdb:
minAvailable: 2
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app.kubernetes.io/name: agenta
app.kubernetes.io/component: api

An empty list removes the constraints for that workload:

api:
topologySpreadConstraints: []

There is no per-workload switch for the budget. To remove every generated object, turn the mechanism off for the whole release:

podDisruptionBudgets:
enabled: false
topologySpread:
enabled: false

The migration Job carries no budget in any case. A budget over it would block the node that runs it.

Keep a rollout from dropping requests​

A managed load balancer keeps sending requests to a pod for a few seconds after Kubernetes removes that pod from the endpoints. A pod that exits as soon as it gets the TERM signal answers those requests with 502.

Three keys per workload close that window. strategy is rendered under the Deployment's spec.strategy. lifecycle is rendered under the container's lifecycle. terminationGracePeriodSeconds is rendered on the pod spec.

These are the chart's defaults. A key you set replaces the default for that workload.

WorkloadstrategyterminationGracePeriodSecondslifecycle.preStop
apiRollingUpdate, maxUnavailable: 0, maxSurge: 160exec sleep 10
servicesRollingUpdate, maxUnavailable: 0, maxSurge: 160exec sleep 10
webRollingUpdate, maxUnavailable: 0, maxSurge: 160exec sleep 10
webMobileRollingUpdate, maxUnavailable: 0, maxSurge: 160exec sleep 10
workerStreamsRollingUpdate, maxUnavailable: 0, maxSurge: 1120none
workerQueuesRollingUpdate, maxUnavailable: 0, maxSurge: 1120none
supertokensRollingUpdate, maxUnavailable: 0, maxSurge: 1nonenone
agentRunnerRecreate300exec sleep 10
cronRecreatenonenone
redisVolatileRecreatenonenone
redisDurablenone (a StatefulSet has no spec.strategy)nonenone
store.seaweedfsnone (a StatefulSet has no spec.strategy)nonenone

Every workload that takes public traffic (api, services, web, webMobile) already has all three, so you do not need to set them. To change one, set the key on that workload. For example, a longer drain for the API:

api:
terminationGracePeriodSeconds: 90

Three constraints apply when you change those values:

  • The grace period must be longer than the preStop delay plus the time the process needs to drain. The KILL signal lands at the end of the grace period.
  • preStop: {sleep: {seconds: 10}} needs Kubernetes 1.30 or later. Below 1.30, use an exec hook. The default exec hook runs sleep directly, so the image needs a sleep binary. The Agenta images have one. A custom image without sleep must set its own lifecycle.
  • maxUnavailable: 0 needs room for one more pod than you run today. On a full cluster the rollout waits for a node instead of starting. From version 0.122.1 this also applies to api.

Lock down the bundled data stores​

By default no NetworkPolicy selects the bundled Redis instances or the bundled SeaweedFS, so any pod in the namespace can reach them. Turn one policy on per bundled store:

networkPolicy:
enabled: true

The chart then renders one NetworkPolicy per bundled store it deploys. Each policy allows ingress only from pods of this release, in the same namespace, on that store's port.

Two things to check before you count a store as locked down:

  • Policies are additive. A namespace-wide allow policy that also selects these pods still lets its own sources in.
  • A cluster whose CNI does not implement NetworkPolicy accepts the objects and ignores them, with no event and no warning. On GKE, turn on network policy enforcement or Dataplane V2.

The bundled PostgreSQL is not covered. The Bitnami subchart renders its own policy, which allows every source on port 5432. Restrict it through the subchart's own values:

postgresql:
primary:
networkPolicy:
allowExternal: false
extraIngress:
- ports:
- port: 5432
from:
- podSelector:
matchLabels:
app.kubernetes.io/instance: '{{ .Release.Name }}'

allowExternal: false narrows the subchart's rule to pods labelled <release>-postgresql-client: "true", which the Agenta workloads do not carry. The extraIngress entry above lets them back in.

Settings that are required on Google Kubernetes Engine​

These are not optional on GKE. Without them the Ingress never gets an address.

Create the address, the certificate and the redirect​

The Ingress below names three objects that the chart does not create: a global static IP address, a Google-managed certificate, and a FrontendConfig that redirects HTTP to HTTPS. Reserve the address once:

gcloud compute addresses create agenta-ip --global

Point the DNS record for your host at that address. Create the other two objects with extraObjects, so the release owns them:

extraObjects:
- apiVersion: networking.gke.io/v1
kind: ManagedCertificate
metadata:
name: agenta-cert
namespace: "{{ .Release.Namespace }}"
spec:
domains:
- agenta.example.com
- apiVersion: networking.gke.io/v1beta1
kind: FrontendConfig
metadata:
name: agenta-https
namespace: "{{ .Release.Namespace }}"
spec:
redirectToHttps:
enabled: true
responseCodeName: MOVED_PERMANENTLY_DEFAULT

GKE provisions the certificate only after the DNS record resolves to the address. This can take 10 to 30 minutes.

Route by annotation, not by class​

The GKE Ingress controller ignores spec.ingressClassName. It reacts only to the legacy kubernetes.io/ingress.class annotation. An Ingress that carries a class name the controller does not own gets no controller events and never gets an IP.

Set ingress.className to an empty string so the chart leaves the field out, then route with the annotation:

agenta:
webUrl: "https://agenta.example.com"
apiUrl: "https://agenta.example.com/api"
servicesUrl: "https://agenta.example.com/services"
ingress:
enabled: true
className: ""
host: agenta.example.com
annotations:
kubernetes.io/ingress.class: gce
kubernetes.io/ingress.global-static-ip-name: agenta-ip
networking.gke.io/managed-certificates: agenta-cert
networking.gke.io/v1beta1.FrontendConfig: agenta-https

An unset ingress.className still defaults to traefik, which is correct for the bundled stack. Only an explicit empty string omits the field.

Set the three public URLs yourself. The chart derives https URLs only when ingress.tls is a non-empty list, and a Google-managed certificate leaves ingress.tls empty. Without them the derived URLs start with http://.

Keep the chart's default paths, which use pathType: Prefix. With ImplementationSpecific, GKE treats /api as an exact path, so a request to /api/... does not reach the API.

The GKE load balancer cannot rewrite a path. It forwards /api and /services verbatim. Both backends accept their own prefix, so nothing has to strip it. See Path handling. Keep the mobile path at /m, because the mobile image is built with that basePath.

Set the NEG annotation on every routed Service​

The GCE Ingress controller reads two annotations off each Service. Set the NEG annotation yourself even on Autopilot, which adds it only when it creates the Service. A helm upgrade --force recreates Services without it.

api:
service:
annotations:
cloud.google.com/neg: '{"ingress": true}'
cloud.google.com/backend-config: '{"default": "agenta-api"}'
web:
service:
annotations:
cloud.google.com/neg: '{"ingress": true}'

webMobile.service.annotations and services.service.annotations work the same way. The BackendConfig objects are yours to create. Put them in extraObjects if you want the release to own them.

GKE Autopilot: turn off the FUSE mount​

A local sandbox mounts the object store prefix with FUSE, so working directories persist across turns. The runner container therefore asks for the SYS_ADMIN capability and a hostPath mount of /dev/fuse whenever store.enabled and agentRunner.fuse.enabled are both true. Autopilot rejects both, and the runner pod never starts.

On Autopilot, set agentRunner.fuse.enabled: false. Local sandboxes still run, but each run gets an ephemeral working directory, and working directories do not persist across turns.

If agents need working directories that persist across turns, run sandboxes on the Daytona provider:

agentRunner:
fuse:
enabled: false
providers:
enabled: [daytona]
default: daytona
daytona:
apiKeySecretRef:
name: agenta-daytona
key: api-key

Run agents in a cloud sandbox (Daytona) covers the Daytona account and the snapshot.

Check a values file before you install​

helm lint hosting/kubernetes/helm -f my-values.yaml
helm template agenta hosting/kubernetes/helm -f my-values.yaml

values.schema.json validates declared fields, and the chart's own validations fail the render with a message that names the key to fix. The root of the values file and the postgresql block accept unknown keys for subchart wiring, so a misspelled key there is accepted and the default applies. Read the rendered output back before you install.

Three checks on the rendered output:

# Save the render once, then read it back.
helm template agenta hosting/kubernetes/helm -f my-values.yaml > rendered.yaml

# Count the budgets. One per workload with two or more replicas.
grep -c 'kind: PodDisruptionBudget' rendered.yaml

# Read the database URI back. A misspelled postgresql key is silently ignored.
grep -A1 'name: POSTGRES_URI_CORE' rendered.yaml

# On GKE, confirm the NEG annotation is on every routed Service.
grep 'cloud.google.com/neg' rendered.yaml