---
title: "AKS: Ship Production-Ready Apps with Helm and CNI Overlay"
description: "Build a production-shaped AKS path: Azure CNI Overlay, ACR, Helm, Gateway API routing, Entra workload identity, and NetworkPolicy hardening."
canonical: "https://adamtheautomator.com/aks-production-ready-helm-overlay/"
---

# AKS: Ship Production-Ready Apps with Helm and CNI Overlay

> Build a production-shaped AKS path: Azure CNI Overlay, ACR, Helm, Gateway API routing, Entra workload identity, and NetworkPolicy hardening.

Source: https://adamtheautomator.com/aks-production-ready-helm-overlay/

---

ATA Learning

Tap to hide

[

ATA Learning

](/)

*   [Home](/)
*   [Tutorials](/tutorials/)
*   [Instructors](/author/)
*   [Advertising](/advertising/)
*   [Recommended Resources](/resources/)
*   [About Adam](/about-adam/)

Search for:  

*   [](https://twitter.com/adbertram)
*   [](https://github.com/Adam-the-Automator)
*   [](https://www.linkedin.com/company/adam-the-automator-llc)
*   [](/feed/)

![AKS: Ship Production-Ready Apps with Helm and CNI Overlay](https://adamtheautomator.com/wp-content/uploads/publisher/3075d9c85b2b8183b1f2c3f391257a46/bd9070a98945da0014fd6ce4e9aeebdec333abbe87cd0491a6f55c3584262ee7.webp)

# AKS: Ship Production-Ready Apps with Helm and CNI Overlay

[![](https://secure.gravatar.com/avatar/d0b9d42e21e5622713f8b693aa5c0f9244d5f7dd200ed29b8398f52dee5de337?s=192&d=mm&r=g)Adam Bertram](https://adamtheautomator.com/author/adam-bertram/)30 September 202615 min. read

Categories: [Cloud](/category/cloud/)

Tags:[Azure](/tag/azure/)[Kubernetes](/tag/kubernetes/)[AKS](/tag/aks/)[Helm](/tag/helm/)

Table of Contents

*   [Prerequisites](#prerequisites)
*   [Confirm Access and Available Tooling](#confirm-access-and-available-tooling)
*   [Verify Your Tooling and Set Resource Names](#verify-your-tooling-and-set-resource-names)
*   [Confirm SKU Availability Before You Provision](#confirm-sku-availability-before-you-provision)
*   [Create a Production-Shaped AKS Cluster](#create-a-production-shaped-aks-cluster)
*   [Confirm Nodes Across Zones](#confirm-nodes-across-zones)
*   [Attach ACR and Confirm Image Pulls](#attach-acr-and-confirm-image-pulls)
*   [Verify or Attach Registry Access](#verify-or-attach-registry-access)
*   [Push a Practice Image to ACR](#push-a-practice-image-to-acr)
*   [Troubleshoot Pull Failures and Gate Releases](#troubleshoot-pull-failures-and-gate-releases)
*   [Deploy the Application with Helm](#deploy-the-application-with-helm)
*   [Scaffold the Chart and Point It at ACR](#scaffold-the-chart-and-point-it-at-acr)
*   [Configure Networking for Production Traffic](#configure-networking-for-production-traffic)
*   [Route Traffic Through the Managed Gateway](#route-traffic-through-the-managed-gateway)
*   [Choose an Exposure Model](#choose-an-exposure-model)
*   [Harden Identity and Network Security](#harden-identity-and-network-security)
*   [Federate a Workload Identity](#federate-a-workload-identity)
*   [Restrict Ingress with NetworkPolicy](#restrict-ingress-with-networkpolicy)
*   [Validate Identity from the Pod](#validate-identity-from-the-pod)
*   [Verify the Release and Capture a Baseline](#verify-the-release-and-capture-a-baseline)
*   [Troubleshoot the Failures You Will Actually See](#troubleshoot-the-failures-you-will-actually-see)
*   [Match the Symptom to Its Cause](#match-the-symptom-to-its-cause)
*   [Work the Stack in Order](#work-the-stack-in-order)
*   [Clean Up Lab Resources](#clean-up-lab-resources)

Don’t copy the AKS quickstart defaults into a subscription that already runs production workloads like your payroll APIs. Those defaults [get a cluster running](https://adamtheautomator.com/azure-kubernetes-service/). But they also skip the zone spread, workload identity, application routing, and network isolation controls that a production cluster needs before it takes real traffic.

You’re going to build a production-shaped [Azure Kubernetes Service (AKS)](https://learn.microsoft.com/en-us/azure/aks/what-is-aks) path end to end. First you create a multi-zone cluster with [Azure CNI Overlay](https://learn.microsoft.com/en-us/azure/aks/concepts-network-azure-cni-overlay), attach [Azure Container Registry (ACR)](https://learn.microsoft.com/en-us/azure/container-registry/container-registry-intro), and ship a sample app with [Helm](https://helm.sh/docs/). Then you publish it through the application routing Gateway API, federate [Microsoft Entra Workload ID](https://learn.microsoft.com/en-us/azure/aks/workload-identity-overview), restrict traffic with [Kubernetes NetworkPolicy](https://kubernetes.io/docs/concepts/services-networking/network-policies/), and record a healthy baseline before you hand the cluster to an application team.

## Prerequisites

Confirm every item below before anything in this walkthrough starts billing your subscription.

### Confirm Access and Available Tooling

If you want to follow along hands-on, gather:

*   An Azure subscription with rights to create resource groups, AKS, ACR, managed identities, and role assignments ([Contributor plus User Access Administrator](https://learn.microsoft.com/en-us/azure/role-based-access-control/built-in-roles), or an equivalent custom role).
    
*   [Azure CLI](https://learn.microsoft.com/en-us/cli/azure/install-azure-cli) 2.86.0 or later (`az version`), signed in with `az login`, plus [kubectl](https://kubernetes.io/docs/reference/kubectl/) from `az aks install-cli`. The AKS-managed [application routing Gateway API](https://learn.microsoft.com/en-us/azure/aks/app-routing-gateway-api) requires Azure CLI 2.86.0 or newer.
    
*   Helm 3.x (`helm version`). If you need it, use the [Helm installation guide](https://helm.sh/docs/v3/intro/install/). Helm 2 is retired; don’t chase old tiller docs.
    
*   A region that supports [AKS availability zones](https://learn.microsoft.com/en-us/azure/aks/availability-zones) for the VM size you pick (eastus2 is a common lab choice).
    
*   Optional: a Dockerfile and app source. The walkthrough uses a public sample image so you can practice the cluster path without waiting on a build pipeline.
    

### Verify Your Tooling and Set Resource Names

Confirm your CLI and Helm versions, then set the shared resource names every later command will reuse:

```bash
az account show --query "{name:name,id:id}" -o table
az version --query '"azure-cli"' -o tsv
helm version --short
```

Those three checks confirm subscription context, CLI currency, and Helm 3 before you create billable resources.

Set shared names once, before the SKU check, so every later command targets the same resources:

```bash
export LOCATION=eastus2
export RG=rg-aks-prod-lab
export CLUSTER=aks-prod-lab
export ACR=acrprodlab$RANDOM
```

### Confirm SKU Availability Before You Provision

Do one more check before provisioning: make sure the VM SKU you plan to use is actually offered in the region and zones you selected. Zone redundancy only works if the subscription can allocate the SKU in every zone you pick. The command below gives you the regional SKU view before AKS turns a quota or placement problem into a long cluster-creation failure:

```bash
az vm list-skus \
  --location "$LOCATION" \
  --size Standard_D4s_v5 \
  --all \
  -o table
```

`--all` keeps sizes your subscription can’t use in the output, so a restricted SKU shows its restriction instead of silently disappearing from the list.

You’re looking for the SKU to be available in the region with no restriction that conflicts with the zones you plan to use. If capacity or quota is tight, change the VM size or region before you create the cluster, while ACR, identities, and networking don’t exist yet.

## Create a Production-Shaped AKS Cluster

A production cluster needs zone spread, a managed identity, overlay networking, and identity federation switches turned on at create time. Microsoft’s own quickstart states that in a production environment, “the recommendation is a node count of three or more nodes,” and `az aks create` defaults to three when you omit `--node-count` ([Deploy an AKS cluster using Azure CLI](https://learn.microsoft.com/en-us/azure/aks/learn/quick-kubernetes-deploy-cli)).

* * *

_**Warning: The next commands create billed Azure resources (VMs, load balancers, ACR storage). Delete the resource group when you finish, or your lab keeps charging overnight.**_

* * *

```bash
az group create --name "$RG" --location "$LOCATION"

az acr create \
  --resource-group "$RG" \
  --name "$ACR" \
  --sku Standard \
  --role-assignment-mode rbac \
  --location "$LOCATION"
```

Standard ACR is enough for this lab. If you later need geo-replication, upgrade the registry to Premium first. This lab pins the registry to RBAC mode because AKS `--attach-acr` uses `AcrPull`, while registries that use [attribute-based access control (ABAC) repository permissions](https://learn.microsoft.com/en-us/azure/container-registry/container-registry-rbac-abac-repository-permissions) require a different repository-reader role and [can’t use the automatic attach flow](https://learn.microsoft.com/en-us/azure/aks/cluster-container-registry-integration). Create the cluster with overlay networking, the [Cilium dataplane](https://learn.microsoft.com/en-us/azure/aks/azure-cni-powered-by-cilium), the [OpenID Connect (OIDC) issuer](https://learn.microsoft.com/en-us/azure/aks/use-oidc-issuer), and workload identity in one pass:

```bash
az aks create \
  --resource-group "$RG" \
  --name "$CLUSTER" \
  --location "$LOCATION" \
  --node-count 3 \
  --node-vm-size Standard_D4s_v5 \
  --zones 1 2 3 \
  --network-plugin azure \
  --network-plugin-mode overlay \
  --network-dataplane cilium \
  --enable-managed-identity \
  --enable-oidc-issuer \
  --enable-workload-identity \
  --enable-gateway-api \
  --enable-app-routing-istio \
  --attach-acr "$ACR" \
  --generate-ssh-keys
```

| Flag | Why it matters here |
| --- | --- |
| `--zones 1 2 3` | Spreads system nodes across zones so a single datacenter blip doesn’t take the pool offline |
| `--network-plugin-mode overlay` | Gives pods a private CIDR instead of consuming a VNet IP per pod |
| `--network-dataplane cilium` | Extended Berkeley Packet Filter (eBPF) dataplane Microsoft recommends with overlay for policy and scale |
| `--enable-oidc-issuer` / `--enable-workload-identity` | Required later for pod-to-Azure auth without secrets in the pod |
| `--enable-gateway-api` / `--enable-app-routing-istio` | Installs the Gateway API custom resource definitions (CRDs) and the `approuting-istio` GatewayClass that back the application routing Gateway API |
| `--attach-acr` | Grants the kubelet identity AcrPull so your first deploy doesn’t stall on ImagePullBackOff |

The diagram below shows how those create-time choices fit together: three nodes across zones, a pod CIDR separate from the node subnet, and a registry the kubelet can pull from.

![AKS production architecture](https://adamtheautomator.com/wp-content/uploads/publisher/35ebbd552a90b1cc2d29b8c5043e3f32e00109282dc89a1ef306764bf97ce4e8.jpg)

### Confirm Nodes Across Zones

When provisioning finishes, pull credentials and confirm three Ready nodes:

```bash
az aks get-credentials --resource-group "$RG" --name "$CLUSTER" --overwrite-existing
kubectl get nodes -o wide
```

You should see three nodes in different availability zones. If a zone column is empty, your region or VM size doesn’t support the zones you requested, and you need to recreate with a supported SKU before you send production traffic to it.

Node spread is only the infrastructure half of availability. But Kubernetes can still place both replicas of an application on the same node or in the same zone unless the workload tells the scheduler what kind of spread you want. For a real service, add [topology spread constraints](https://kubernetes.io/docs/concepts/scheduling-eviction/topology-spread-constraints/) or pod anti-affinity to the chart and verify the result with `kubectl get pods -o wide`. Losing one node or one zone shouldn’t remove every ready replica of the application.

Treat node distribution and pod distribution as two separate checks.

## Attach ACR and Confirm Image Pulls

You should verify the registry path even when you created the cluster with `--attach-acr`.

### Verify or Attach Registry Access

A missing registry grant surfaces as a Helm or Kubernetes error, so it sends you debugging the wrong layer. Microsoft documents the attach path in [Authenticate with ACR from AKS](https://learn.microsoft.com/en-us/azure/aks/cluster-container-registry-integration).

```bash
az aks check-acr \
  --resource-group "$RG" \
  --name "$CLUSTER" \
  --acr "${ACR}.azurecr.io"
```

A successful check prints that the registry is reachable from a node. If you created the cluster without `--attach-acr`, attach it now:

```bash
az aks update \
  --resource-group "$RG" \
  --name "$CLUSTER" \
  --attach-acr "$ACR"
```

### Push a Practice Image to ACR

Push a practice image into the registry so Helm has something private to pull. [ACR’s server-side import](https://learn.microsoft.com/en-us/azure/container-registry/container-registry-import-images) keeps Docker off your laptop:

```bash
az acr import \
  --name "$ACR" \
  --source mcr.microsoft.com/azuredocs/aci-helloworld:latest \
  --image demo/web:1.0.0
```

### Troubleshoot Pull Failures and Gate Releases

The next table maps common pull failures to their most likely cause.

| Symptom | Likely cause | First check |
| --- | --- | --- |
| `ImagePullBackOff` | Missing AcrPull on kubelet identity | `az aks check-acr` |
| `ErrImagePull` with 401 | Wrong login server in the manifest | Compare `image:` to `az acr show -n $ACR --query loginServer` |
| Pull works in one pool only | Kubelet identity missing from that pool’s VM scale set | Compare `identityProfile.kubeletidentity` with the VM scale set identity |

Keep that table in your on-call runbook. When a ticket says Helm is broken, check registry RBAC before you touch the chart.

Make the registry check a release gate instead of a troubleshooting step you remember after a failed deployment. Capture the login server once, confirm the image exists, and inspect the digest that ACR reports:

```bash
ACR_LOGIN_SERVER=$(az acr show -n "$ACR" --query loginServer -o tsv)
az acr repository show-tags -n "$ACR" --repository demo/web -o table
az acr manifest list-metadata --registry "$ACR" --name demo/web -o table
```

Helm requests a tag, and ACR resolves it to a digest. Pin the application version to an immutable tag or digest in your release process so the same tag never points at different image contents.

## Deploy the Application with Helm

[Helm on AKS](https://learn.microsoft.com/en-us/azure/aks/kubernetes-helm) packages environment-specific values into a release with a stable name and a revision history you can roll back to.

### Scaffold the Chart and Point It at ACR

Scaffold a chart, then point it at the ACR image:

```bash
helm create demo-web
```

Edit `demo-web/values.yaml` so the image and service match the registry you just filled:

```yaml
replicaCount: 2

image:
  repository: <ACR_LOGIN_SERVER>/demo/web
  pullPolicy: IfNotPresent
  tag: "1.0.0"

service:
  type: ClusterIP
  port: 80

resources:
  requests:
    cpu: 50m
    memory: 64Mi
  limits:
    cpu: 250m
    memory: 256Mi
```

Replace `<ACR_LOGIN_SERVER>` with the output of `az acr show --name "$ACR" --query loginServer -o tsv`. Keep `service.type` as ClusterIP. You configure the shared Gateway that fronts it later in this walkthrough.

`replicaCount: 2` gives the scheduler two pods to spread, which the topology spread check earlier depends on. `resources.requests` is what the scheduler reserves on a node when it places each pod, and `resources.limits` caps the container: CPU above the limit gets throttled, and memory above it gets the container killed.

Install with an atomic upgrade so a bad chart revision rolls back instead of leaving a half-applied release:

```bash
kubectl create namespace demo

helm upgrade --install demo-web ./demo-web \
  --namespace demo \
  --atomic \
  --timeout 5m \
  --set image.repository="${ACR_LOGIN_SERVER}/demo/web" \
  --set image.tag=1.0.0
```

Helm’s [`--atomic` upgrade flag](https://helm.sh/docs/v3/helm/helm_upgrade/) waits for resources to become ready and reverts on failure. Without it, a bad tag leaves CrashLoopBackOff pods in the namespace until you roll back yourself.

After Helm exits, inspect the values it stored, the resources Kubernetes accepted, and the current revision number:

```bash
helm get values demo-web -n demo
helm status demo-web -n demo
helm history demo-web -n demo
```

Those commands answer different questions. Start with `helm status` to see the release state and current revision. Use `helm get values` when you need the configuration that actually landed, and check `helm history` when you need a rollback target. Save those outputs in CI when you promote the same chart between environments.

Then confirm Kubernetes agrees with Helm about what is running:

```bash
helm list -n demo
kubectl get pods,svc -n demo
kubectl rollout status deployment/demo-web -n demo
```

`helm list` should show the release as `deployed`, `kubectl get pods,svc` should show two Running pods behind a ClusterIP Service, and `kubectl rollout status` returns only when every replica is available. If rollout status hangs, describe the pods before you rerun Helm.

* * *

_**Pro Tip: Pin chart and app versions in CI. Floating _**`latest`**_ tags make _**`helm rollback`**_ lie to you because the tag moved under the old revision.**_

* * *

## Configure Networking for Production Traffic

Exposing application traffic is where quickstart defaults and production diverge. Azure CNI Overlay keeps node IPs in the VNet while pods use a separate CIDR, so pods scale without consuming a VNet address each. Microsoft Learn’s Azure CNI Overlay overview states: “In overlay networking, only the Kubernetes cluster nodes are assigned IPs from subnets. Pods receive IPs from a private Classless Inter-Domain Routing (CIDR) range provided at the time of cluster creation.” ([Overview of Azure CNI Overlay networking](https://learn.microsoft.com/en-us/azure/aks/concepts-network-azure-cni-overlay))

OSS ingress-nginx is not a safe choice for a new production AKS cluster. The Kubernetes project [retired Ingress NGINX in March 2026](https://kubernetes.io/blog/2025/11/11/ingress-nginx-retirement/), meaning no new releases, bug fixes, or security fixes. [AKS supports the application routing Gateway API](https://learn.microsoft.com/en-us/azure/aks/app-routing-gateway-api) for new workloads, and you already enabled its Gateway API CRDs and `approuting-istio` implementation when you created the cluster.

### Route Traffic Through the Managed Gateway

Verify the managed control plane and GatewayClass before you publish any route:

```bash
kubectl get pods -n aks-istio-system
kubectl get gatewayclass approuting-istio
```

Create a `Gateway` and `HTTPRoute` that send `/` traffic to the Service created by your Helm release:

```yaml
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: demo-web-gateway
  namespace: demo
spec:
  gatewayClassName: approuting-istio
  listeners:
    - name: http
      port: 80
      protocol: HTTP
      allowedRoutes:
        namespaces:
          from: Same
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: demo-web
  namespace: demo
spec:
  parentRefs:
    - name: demo-web-gateway
  rules:
    - matches:
        - path:
            type: PathPrefix
            value: /
      backendRefs:
        - name: demo-web
          port: 80
```

```bash
kubectl apply -f demo-gateway.yaml

kubectl wait -n demo --for=condition=programmed gateway/demo-web-gateway --timeout=300s

kubectl get gateway,httproute -n demo
```

The `Programmed=True` condition is the point where the managed controller has provisioned the Gateway data plane. The application routing Gateway API also creates a Deployment, Service, HorizontalPodAutoscaler (HPA), and PodDisruptionBudget for each Gateway. Verify that generated infrastructure directly:

```bash
kubectl get deployment,service,hpa,pdb -n demo -l gateway.networking.k8s.io/gateway-name=demo-web-gateway
```

### Choose an Exposure Model

| Exposure choice | Use when | Cost / risk note |
| --- | --- | --- |
| `ClusterIP` only | East-west or internal callers | No public IP; pair with private connectivity when the caller is outside the cluster |
| Public `LoadBalancer` Service | Break-glass demos or a single simple TCP/UDP service | One public IP per Service; easy to forget and keep paying for |
| AKS application routing Gateway API | Managed HTTP/HTTPS host and path routing | AKS-managed infrastructure with a supported upgrade path |

Pick the exposure model before you attach DNS and certificates. If the service is internal, keep it private and prove that the caller path works from the networks that need it. If it’s internet-facing, treat the Gateway as shared infrastructure: define who owns DNS, TLS, public IP changes, policy, and rollback. An application chart owner shouldn’t be able to turn a private service public just by changing `service.type`.

The comparison below shows the three exposure models side by side, including where each one puts the public IP.

![AKS exposure options](https://adamtheautomator.com/wp-content/uploads/publisher/c184d71339b785b718af4aa4883ee75bd8c77d4ffea6096488439d618f43ece9.jpg)

Also verify the path from both directions. From outside the cluster, confirm the Gateway address and HTTP status. From inside the cluster, hit the `ClusterIP` Service directly. If the internal request works but the Gateway path fails, the problem sits in Gateway programming, the `HTTPRoute`, the managed proxy, or the Azure load balancer.

## Harden Identity and Network Security

Production AKS expects [Workload ID for Azure resource access](https://adamtheautomator.com/implementing-workload-identity-aks/) and NetworkPolicy so a compromised pod can’t reach every other pod on the cluster’s flat network. Microsoft Learn’s [Workload ID overview](https://learn.microsoft.com/en-us/azure/aks/workload-identity-overview) describes the integration this way: “Microsoft Entra Workload ID integrates with the capabilities native to Kubernetes to federate with external identity providers, allowing you to assign workload identities to your workloads to authenticate and access other services and resources.”

### Federate a Workload Identity

You already enabled the OIDC issuer and workload identity at cluster create. Wire a sample identity for the `demo` namespace:

```bash
export UAMI=uami-demo-web
az identity create --name "$UAMI" --resource-group "$RG" --location "$LOCATION"
export CLIENT_ID=$(az identity show -g "$RG" -n "$UAMI" --query clientId -o tsv)
export OIDC_ISSUER=$(az aks show -g "$RG" -n "$CLUSTER" --query oidcIssuerProfile.issuerUrl -o tsv)
```

```bash
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: ServiceAccount
metadata:
  name: demo-web
  namespace: demo
  annotations:
    azure.workload.identity/client-id: ${CLIENT_ID}
EOF
```

```bash
az identity federated-credential create \
  --name demo-web-fed \
  --identity-name "$UAMI" \
  --resource-group "$RG" \
  --issuer "$OIDC_ISSUER" \
  --subject "system:serviceaccount:demo:demo-web" \
  --audience api://AzureADTokenExchange
```

`--issuer` is the cluster’s OIDC issuer URL you captured above, so Microsoft Entra ID accepts tokens only from this cluster. `--subject` must exactly match `system:serviceaccount:<namespace>:<service-account>`, which is `demo:demo-web` here; if it differs by one character, Entra ID refuses to issue the token. `--audience` must match the `aud` claim in the projected token, and `api://AzureADTokenExchange` is the value Workload ID expects.

Label pods that should receive the projected token with `azure.workload.identity/use: "true"` and set `serviceAccountName: demo-web` in the Deployment (or chart values). Without that label, the mutating webhook skips the pod and your app falls back to whatever local credential mistake someone left in the image.

### Restrict Ingress with NetworkPolicy

Then isolate ingress to the `demo-web` pods with [Kubernetes NetworkPolicy](https://adamtheautomator.com/practical-guide-kubernetes-security-management/). This first policy doesn’t default-deny the whole namespace, and it deliberately leaves egress unrestricted because Workload ID needs outbound HTTPS to Microsoft Entra. The cluster uses Cilium, so the standard Kubernetes L3/L4 policy below is enforced:

```yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: demo-web-ingress
  namespace: demo
spec:
  podSelector:
    matchLabels:
      app.kubernetes.io/name: demo-web
  policyTypes:
    - Ingress
  ingress:
    - from:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: demo
      ports:
        - protocol: TCP
          port: 80
```

For this lab, the ingress rule allows port 80 from the `demo` namespace because the AKS-managed Gateway proxy for `demo-web-gateway` is provisioned in that namespace. In a long-lived environment, inspect the generated Gateway proxy labels and add a `podSelector` so the policy permits only that data plane rather than every pod in `demo`.

If you later restrict egress, account for the [AKS outbound requirements](https://learn.microsoft.com/en-us/azure/aks/outbound-rules-control-egress) instead of copying an IP-only allowlist. Workload ID requires HTTPS 443 to `login.microsoftonline.com`, while standard Kubernetes NetworkPolicy selects pods, namespaces, and CIDR blocks rather than FQDNs. On Cilium, [Advanced Container Networking Services](https://learn.microsoft.com/en-us/azure/aks/advanced-container-networking-services-overview) can add FQDN filtering if you enable it; otherwise enforce FQDN egress at a firewall or another egress layer.

### Validate Identity from the Pod

Validate identity from the pod that will use it. A wrong service account, pod label, or federated subject breaks the token exchange even when every Azure object exists. Start by confirming what the pod actually received:

```bash
kubectl get serviceaccount demo-web -n demo -o yaml
kubectl get pod -n demo -l app.kubernetes.io/name=demo-web -o yaml
```

Check for the expected service account and workload-identity label before you troubleshoot Azure RBAC. Then use the application itself, or a temporary diagnostic pod that uses the same service account, to request the Azure resource it needs. That separates federation problems from permission problems: no projected identity points you back to Kubernetes configuration, while a valid token followed by an authorization failure points you toward the Azure role assignment.

* * *

_**Reality Check: A new federated credential can take a few seconds to propagate (**_[_**Workload ID deployment guide**_](https://learn.microsoft.com/en-us/azure/aks/workload-identity-deploy-cluster)_**). If the first token exchange fails, wait a minute and retry before rewriting the subject string.**_

* * *

## Verify the Release and Capture a Baseline

Your deployment includes verification. Wait for the Gateway, pull its address, hit the path, and record pod readiness so the next change has a baseline. The [application routing Gateway API guidance](https://learn.microsoft.com/en-us/azure/aks/app-routing-gateway-api) follows the same sequence; keep that habit when Helm owns the backend release.

```bash
GATEWAY_IP=$(kubectl get gateway demo-web-gateway -n demo -o jsonpath='{.status.addresses[0].value}')

echo "$GATEWAY_IP"
curl -sS -o /dev/null -w "%{http_code}\n" "http://${GATEWAY_IP}/"

kubectl get pods -n demo -o wide
kubectl describe gateway demo-web-gateway -n demo
kubectl describe httproute demo-web -n demo
```

Expect HTTP 200 from the sample image. If you get 502 or 503, the Gateway proxy can’t reach ready endpoints. Check endpoints before you touch DNS:

```bash
kubectl get endpoints demo-web -n demo
kubectl logs -n demo deployment/demo-web-gateway-approuting-istio --tail=50
```

Record the revision number that `helm history` marks as `deployed`, because `helm rollback demo-web <revision> -n demo` needs it when the next change breaks:

```bash
helm history demo-web -n demo
```

Capture a small handoff bundle while the release is healthy, just enough state to tell whether tomorrow’s failure is new:

```bash
helm status demo-web -n demo
kubectl get deploy,pods,svc,gateway,httproute -n demo -o wide
kubectl get events -n demo --sort-by=.lastTimestamp | tail -n 20
```

Record the Helm revision, image tag, ready replica count, Gateway address, `HTTPRoute` status, and the HTTP result you just tested. When a later rollout fails, compare those few artifacts first. If the image and values are unchanged but the Gateway address or route conditions changed, investigate networking. If the Gateway is stable but ready replicas dropped after a new revision, investigate the workload.

## Troubleshoot the Failures You Will Actually See

Work the list below in order, and keep the [ACR integration](https://learn.microsoft.com/en-us/azure/aks/cluster-container-registry-integration), [Gateway API routing](https://learn.microsoft.com/en-us/azure/aks/app-routing-gateway-api), and [network policy](https://learn.microsoft.com/en-us/azure/aks/use-network-policies) pages open as your operator reference.

### Match the Symptom to Its Cause

Each numbered item pairs a symptom with the command that confirms its cause.

1.  **ImagePullBackOff:** Run `az aks check-acr`. Confirm the image string uses `yourregistry.azurecr.io/repo:tag` with no typo in the login server.
    
2.  **Gateway never reaches **`Programmed=True`**:** Run `kubectl describe gateway demo-web-gateway -n demo`, confirm `kubectl get gatewayclass approuting-istio` reports the managed class, and check the `aks-istio-system` control-plane pods. If the Gateway has no address after programming, inspect Azure public IP quota and load balancer events.
    
3.  **CrashLoopBackOff after Helm upgrade:** `helm history` plus `kubectl logs`. If you used `--atomic`, confirm whether Helm already rolled back while you were still staring at the old ReplicaSet.
    
4.  **Workload identity 401 to Key Vault or Storage:** Federated subject must match `system:serviceaccount:<namespace>:<sa>`. Pod label `azure.workload.identity/use=true` must be present. Role assignment must land on the user-assigned identity, not the cluster identity.
    
5.  **Pods Ready but the Gateway returns 404:** The `HTTPRoute` parent reference, path match, backend Service name, or backend port is wrong. `kubectl describe httproute demo-web -n demo` shows whether the route was accepted and which backend it resolved.
    

### Work the Stack in Order

Work the stack from the resource boundary inward instead of changing three things at once. First confirm the Azure resources exist and the cluster is reachable. Then check node health, pod scheduling, image pulls, Service endpoints, Gateway programming, and finally route status. Stop at the first layer that disagrees with the healthy baseline.

```bash
az aks show -g "$RG" -n "$CLUSTER" --query provisioningState -o tsv
kubectl get nodes
kubectl get pods -n demo -o wide
kubectl get endpoints demo-web -n demo
kubectl describe gateway demo-web-gateway -n demo
kubectl describe httproute demo-web -n demo
kubectl get events -n demo --sort-by=.lastTimestamp | tail -n 20
```

Events tell you whether the failure is scheduling, pulling, probing, or provisioning. The endpoint check tells you whether the Service has healthy backends before traffic ever reaches the managed Gateway proxy. Checking endpoints first stops you from editing route configuration when the real problem is an image that never started.

## Clean Up Lab Resources

You will pay for a lab cluster that keeps running after you stop using it. When you’re done practicing, remove the application resources and then delete the resource group, matching the cleanup guidance in Microsoft’s [AKS CLI quickstart](https://learn.microsoft.com/en-us/azure/aks/learn/quick-kubernetes-deploy-cli).

```bash
helm uninstall demo-web -n demo
kubectl delete httproute demo-web -n demo --ignore-not-found
kubectl delete gateway demo-web-gateway -n demo --ignore-not-found
az group delete --name "$RG" --yes --no-wait
```

`--no-wait` returns your shell while Azure deletes in the background. Confirm later with `az group show --name "$RG"` until it returns not found.

If this lab is becoming a permanent shared environment, skip deletion and plan the next pass instead: [Azure Policy for AKS](https://learn.microsoft.com/en-us/azure/aks/use-azure-policy), [Microsoft Defender for Containers](https://learn.microsoft.com/en-us/azure/defender-for-cloud/defender-for-containers-introduction), a [private API server](https://learn.microsoft.com/en-us/azure/aks/private-clusters), managed DNS and TLS, and the rest of your ingress hardening. You already run the application routing Gateway API for standard HTTP/HTTPS routing. If you need features outside its limits, evaluate [Application Gateway for Containers](https://learn.microsoft.com/en-us/azure/application-gateway/for-containers/overview) or the [Istio service-mesh Gateway API path](https://learn.microsoft.com/en-us/azure/aks/istio-gateway-api) instead of falling back to retired OSS ingress-nginx.

Before an enterprise app team gets a namespace on this cluster, rerun the baseline you captured: nodes populated across zones, `az aks check-acr` passing, the Gateway reporting `Programmed=True`, the NetworkPolicy in place, and a token exchange working from the `demo-web` pod.

Share this article

[Share on X](https://twitter.com/intent/tweet?url=https%3A%2F%2Fadamtheautomator.com%2Faks-production-ready-helm-overlay%2F&text=AKS%3A%20Ship%20Production-Ready%20Apps%20with%20Helm%20and%20CNI%20Overlay)[Share on Facebook](https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fadamtheautomator.com%2Faks-production-ready-helm-overlay%2F)[Share on LinkedIn](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fadamtheautomator.com%2Faks-production-ready-helm-overlay%2F)

## Related Posts

![](https://adamtheautomator.com/wp-content/uploads/2026/06/featured_image-22.png)

### [Implement Workload Identity in AKS](/implement-workload-identity-aks/)

A deep dive into replacing vulnerable service account credentials with Microsoft Entra Workload Identity in Azure Kubernetes Service (AKS). The post covers federated identity configuration, managed

![](https://adamtheautomator.com/wp-content/uploads/2026/06/26997-implementing-workload-identity-aks-codex.webp)

### [Implementing Workload Identity in AKS](/implementing-workload-identity-aks/)

Implement Microsoft Entra Workload ID in AKS to replace static pod credentials with federated identity and managed identity access.

![](https://adamtheautomator.com/wp-content/uploads/2026/04/featured_image-9.webp)

### [Entra Workload Identity on AKS: No More Secrets](/entra-workload-identity-aks-no-more-secrets/)

Learn how to eliminate Kubernetes secrets by configuring Entra Workload Identity on AKS using OIDC federation, with Bicep and Terraform IaC examples.

## Categories

*   [IT Ops](/category/it-ops/)
*   [Cloud](/category/cloud/)
*   [DevOps](/category/devops/)
*   [Home Ops](/category/home-ops/)
*   [Information Security](/category/infosec/)
*   [Software Development](/category/software-development/)

## Site

*   [Home](/)
*   [Tutorials](/tutorials/)
*   [Instructors](/author/)
*   [Advertising](/advertising/)
*   [Recommended Resources](/resources/)
*   [About Adam](/about-adam/)

Copyright 2026© ATA Learning | [Privacy Policy](/privacy/)
