Configure Istio multi-cluster disaster recovery and automated failover

Advanced 90 min Oct 05, 2026 47 views
Ubuntu 24.04 Debian 12 AlmaLinux 9 Rocky Linux 9

Build a production-grade Istio primary-remote multi-cluster topology with shared trust, locality-aware failover, and automated recovery testing. Covers cross-cluster service discovery, Kiali and Prometheus monitoring, and control plane backup procedures.

Prerequisites

  • Two Kubernetes clusters (1.28+) with routable pod networks and no overlapping CIDRs
  • kubectl configured with contexts for both clusters
  • istioctl 1.23 or newer
  • OpenSSL for CA certificate generation
  • GPG for encrypted backup storage

What this solves

A single-cluster Istio mesh is a single point of failure. This tutorial configures a primary-remote multi-cluster topology spanning two Kubernetes clusters with shared trust, cross-cluster service discovery and locality-aware load balancing so traffic automatically fails over when a cluster or region goes down.

You will build the full trust chain, test real outage scenarios, wire up monitoring with Kiali and Prometheus, and establish backup and recovery procedures for the control plane configuration itself.

Note: This tutorial assumes two existing Kubernetes clusters (cluster-east and cluster-west) with network connectivity between pod CIDRs, either via VPC peering, a flat network, or a VPN mesh such as WireGuard multi-site mesh networking.

Prerequisites and multi-cluster architecture overview

Before starting, confirm each cluster meets the baseline requirements. Both clusters need direct pod-to-pod IP routing (no overlapping CIDRs), a shared root CA for mTLS trust, and API server access from your workstation via distinct kubeconfig contexts.

RequirementDetails
Kubernetes version1.28 or newer on both clusters
Istio version1.22 or newer (istioctl and control plane must match)
NetworkNon-overlapping pod/service CIDRs, routable between clusters
DNSCluster API servers resolvable from each other for cross-cluster secrets
CAShared intermediate CA or common root for mesh-wide mTLS trust

The topology used here is primary-remote: cluster-east runs a full Istio control plane (istiod), and cluster-west runs a remote configuration that uses cluster-east's istiod for configuration while running its own data plane. This is simpler to operate than primary-primary and is sufficient for most disaster recovery requirements.

Install istioctl and verify cluster access

Install the Istio CLI and confirm both kubeconfig contexts are reachable before making any changes.

curl -L https://istio.io/downloadIstio | ISTIO_VERSION=1.23.2 sh -
sudo mv istio-1.23.2/bin/istioctl /usr/local/bin/
istioctl version --remote=false
kubectl config get-contexts
kubectl --context=cluster-east get nodes
kubectl --context=cluster-west get nodes

Setting up primary-remote multi-cluster topology with shared trust domain

Generate a shared root CA

Every cluster in the mesh must trust the same root certificate authority so cross-cluster mTLS connections succeed. Generate an intermediate CA per cluster signed by one shared root.

mkdir -p ~/istio-ca/certs && cd ~/istio-ca/certs
curl -sL https://raw.githubusercontent.com/istio/istio/release-1.23/tools/certs/Makefile.selfsigned.mk -o Makefile
make -f Makefile root-ca
make -f Makefile cluster-east-cacerts
make -f Makefile cluster-west-cacerts

Install the CA secrets into each cluster

The istio-system namespace must exist before the cacerts secret is created, since istiod reads this secret at startup to establish the mesh trust domain.

kubectl --context=cluster-east create namespace istio-system
kubectl --context=cluster-east create secret generic cacerts -n istio-system \
  --from-file=cluster-east/ca-cert.pem \
  --from-file=cluster-east/ca-key.pem \
  --from-file=cluster-east/root-cert.pem \
  --from-file=cluster-east/cert-chain.pem
kubectl --context=cluster-west create namespace istio-system
kubectl --context=cluster-west create secret generic cacerts -n istio-system \
  --from-file=cluster-west/ca-cert.pem \
  --from-file=cluster-west/ca-key.pem \
  --from-file=cluster-west/root-cert.pem \
  --from-file=cluster-west/cert-chain.pem

Label clusters with network and topology metadata

Istio uses the topology.istio.io/network label to determine which endpoints require the east-west gateway versus direct pod routing.

kubectl --context=cluster-east label namespace istio-system topology.istio.io/network=network-east
kubectl --context=cluster-west label namespace istio-system topology.istio.io/network=network-west

Install the primary cluster control plane

Cluster-east becomes the primary with a full istiod deployment, tagged with its mesh, network and cluster identifiers.

apiVersion: install.istio.io/v1alpha1
kind: IstioOperator
spec:
  values:
    global:
      meshID: mesh1
      multiCluster:
        clusterName: cluster-east
      network: network-east
istioctl install --context=cluster-east -f /tmp/cluster-east.yaml -y

Deploy the east-west gateway on the primary

The east-west gateway exposes istiod and cross-cluster service endpoints over mTLS so cluster-west can reach services in cluster-east.

cd ~/istio-1.23.2
samples/multicluster/gen-eastwest-gateway.sh --mesh mesh1 --cluster cluster-east --network network-east | \
  istioctl --context=cluster-east install -y -f -
kubectl --context=cluster-east apply -n istio-system -f samples/multicluster/expose-istiod.yaml
kubectl --context=cluster-east apply -n istio-system -f samples/multicluster/expose-services.yaml

Install the remote cluster configuration

Cluster-west gets the east-west gateway plus a remote config that points it at cluster-east's exposed istiod service.

samples/multicluster/gen-eastwest-gateway.sh --mesh mesh1 --cluster cluster-west --network network-west | \
  istioctl --context=cluster-west install -y -f -
apiVersion: install.istio.io/v1alpha1
kind: IstioOperator
spec:
  profile: remote
  values:
    istiodRemote:
      injectionPath: /inject/cluster/cluster-west/net/network-west
    global:
      remotePilotAddress: 203.0.113.10
      meshID: mesh1
      multiCluster:
        clusterName: cluster-west
      network: network-west
istioctl install --context=cluster-west -f /tmp/cluster-west.yaml -y
Warning: Replace 203.0.113.10 with the actual external IP of the cluster-east east-west gateway LoadBalancer service. Using the wrong address breaks remote config sync silently.

Install a cross-cluster remote secret

Cluster-east's istiod needs a kubeconfig secret for cluster-west so it can watch remote endpoints and services.

istioctl create-remote-secret --context=cluster-west --name=cluster-west | \
  kubectl apply -f - --context=cluster-east

Configuring cross-cluster service discovery and endpoint synchronization

With the remote secret installed, istiod on cluster-east aggregates endpoints from both clusters into a single service registry. Any workload with a matching Service name and namespace across clusters is automatically treated as a single logical service with endpoints in both locations.

Deploy a sample multi-cluster service

Deploy identical Deployment and Service manifests to both clusters. Istio merges the endpoints transparently.

apiVersion: v1
kind: Service
metadata:
  name: helloworld
  labels:
    app: helloworld
spec:
  ports:
  - port: 5000
    name: http
  selector:
    app: helloworld
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: helloworld-v1
spec:
  replicas: 1
  selector:
    matchLabels:
      app: helloworld
  template:
    metadata:
      labels:
        app: helloworld
        version: v1
    spec:
      containers:
      - name: helloworld
        image: docker.io/istio/examples-helloworld-v1
        ports:
        - containerPort: 5000
kubectl --context=cluster-east create namespace sample
kubectl --context=cluster-east label namespace sample istio-injection=enabled
kubectl --context=cluster-east apply -n sample -f /tmp/helloworld.yaml

kubectl --context=cluster-west create namespace sample
kubectl --context=cluster-west label namespace sample istio-injection=enabled
kubectl --context=cluster-west apply -n sample -f /tmp/helloworld.yaml

Confirm endpoint synchronization

Query istiod's debug endpoint to confirm it sees endpoints from both clusters for the same service.

istioctl --context=cluster-east proxy-config endpoint deploy/helloworld-v1.sample -n sample | grep 5000

Implementing locality-aware load balancing and failover policies

Istio's locality load balancing prefers endpoints in the same region and zone as the client, falling back to remote clusters only when local endpoints become unhealthy. This requires outlier detection to be enabled and locality labels to be set correctly on nodes.

Verify node locality labels

Istio reads region and zone from standard Kubernetes topology labels on each node.

kubectl --context=cluster-east get nodes -L topology.kubernetes.io/region,topology.kubernetes.io/zone
kubectl --context=cluster-west get nodes -L topology.kubernetes.io/region,topology.kubernetes.io/zone

Create a DestinationRule with locality failover

This policy sends traffic to the local cluster's region first and automatically fails over to the remote region if local endpoints are ejected by outlier detection.

apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
  name: helloworld
  namespace: sample
spec:
  host: helloworld.sample.svc.cluster.local
  trafficPolicy:
    connectionPool:
      tcp:
        maxConnections: 100
    outlierDetection:
      consecutive5xxErrors: 3
      interval: 10s
      baseEjectionTime: 30s
      maxEjectionPercent: 100
    loadBalancer:
      localityLbSetting:
        enabled: true
        failover:
        - from: eu-west-1
          to: eu-central-1
kubectl --context=cluster-east apply -f /tmp/helloworld-dr.yaml
kubectl --context=cluster-west apply -f /tmp/helloworld-dr.yaml
Note: outlierDetection is mandatory for failover to trigger. Without it, Istio has no signal that local endpoints are unhealthy and will keep sending traffic to a dead cluster.

Set a failover priority with ServiceEntry (optional for external services)

If part of your mesh talks to external endpoints, define locality priority there too so DR behavior is consistent mesh-wide.

kubectl --context=cluster-east get destinationrule helloworld -n sample -o yaml

This pattern complements traffic-splitting strategies covered in Istio multi-cluster canary deployments and pairs well with the resilience patterns in Istio circuit breaker and retry policies.

Testing automated failover scenarios and cluster outage simulation

Generate baseline traffic

Use a simple loop from a client pod to confirm traffic normally stays local to cluster-east.

kubectl --context=cluster-east run sleep --image=curlimages/curl -n sample -- sleep infinity
kubectl --context=cluster-east exec -n sample sleep -- sh -c "for i in $(seq 1 20); do curl -s helloworld.sample:5000/hello; done"

Simulate a cluster-east outage

Scale the local deployment to zero to simulate cluster-east becoming unavailable, then confirm traffic fails over to cluster-west.

kubectl --context=cluster-east scale deploy helloworld-v1 -n sample --replicas=0
kubectl --context=cluster-east exec -n sample sleep -- sh -c "for i in $(seq 1 20); do curl -s helloworld.sample:5000/hello; done"

Responses should now come from version v1 running in cluster-west. Restore the deployment once the test is complete.

kubectl --context=cluster-east scale deploy helloworld-v1 -n sample --replicas=1

Simulate a full network partition

A harder test is severing the east-west gateway connection itself rather than just the workload, which exercises the outlier detection path end to end.

kubectl --context=cluster-east scale deploy istio-eastwestgateway -n istio-system --replicas=0
kubectl --context=cluster-east exec -n sample sleep -- sh -c "for i in $(seq 1 20); do curl -s -w ' -> %{http_code}\n' helloworld.sample:5000/hello; done"
kubectl --context=cluster-east scale deploy istio-eastwestgateway -n istio-system --replicas=1

Monitoring multi-cluster health with Kiali and Prometheus

Install Kiali and Prometheus addons on the primary

Kiali visualizes cross-cluster traffic flow, while Prometheus scrapes per-cluster mesh metrics for alerting.

kubectl --context=cluster-east apply -f samples/addons/prometheus.yaml
kubectl --context=cluster-east apply -f samples/addons/kiali.yaml
kubectl --context=cluster-east rollout status deployment/kiali -n istio-system

Configure remote cluster metrics scraping

Add cluster-west as a remote Prometheus target so dashboards show both clusters in one pane.

kubectl --context=cluster-west get svc istiod -n istio-system
kubectl --context=cluster-east port-forward svc/kiali 20001:20001 -n istio-system

Open the Kiali dashboard and confirm both cluster-east and cluster-west appear under the mesh graph with cross-cluster edges between helloworld instances.

For alerting on failover events and endpoint health degradation, extend this setup with Prometheus federation for multi-cluster monitoring and route notifications through Alertmanager webhook integrations.

Backup and recovery procedures for Istio control plane configuration

Back up IstioOperator manifests and CRDs

The control plane's desired state lives in your IstioOperator YAML and any custom Istio resources (VirtualService, DestinationRule, Gateway). Export all of it regularly.

mkdir -p ~/istio-backup/$(date +%F)
cd ~/istio-backup/$(date +%F)
kubectl --context=cluster-east get istiooperator -A -o yaml > istio-operator-east.yaml
kubectl --context=cluster-east get virtualservices,destinationrules,gateways,serviceentries -A -o yaml > istio-resources-east.yaml
kubectl --context=cluster-west get virtualservices,destinationrules,gateways,serviceentries -A -o yaml > istio-resources-west.yaml

Back up CA material securely

The cacerts secret is the trust root for the entire mesh. Losing it means every workload certificate must be reissued. Export it to an encrypted archive, never in plain text.

kubectl --context=cluster-east get secret cacerts -n istio-system -o yaml > cacerts-east.yaml
gpg --symmetric --cipher-algo AES256 cacerts-east.yaml
rm cacerts-east.yaml

Store the resulting cacerts-east.yaml.gpg in the same offsite location used for your other encrypted secrets, following the key rotation practices in backup encryption key rotation and management.

Automate the backup with a systemd timer

Schedule the export to run daily so configuration drift never outpaces your last known-good backup.

[Unit]
Description=Istio control plane configuration backup

[Service]
Type=oneshot
ExecStart=/usr/local/bin/istio-backup.sh
[Unit]
Description=Daily Istio config backup

[Timer]
OnCalendar=daily
Persistent=true

[Install]
WantedBy=timers.target
sudo systemctl daemon-reload
sudo systemctl enable --now istio-backup.timer

Restore procedure

To rebuild a destroyed cluster, reinstall cacerts first, then reapply the IstioOperator manifest, then reapply custom resources.

gpg --decrypt cacerts-east.yaml.gpg > cacerts-east.yaml
kubectl --context=cluster-east apply -f cacerts-east.yaml
istioctl install --context=cluster-east -f istio-operator-east.yaml -y
kubectl --context=cluster-east apply -f istio-resources-east.yaml

Troubleshooting split-brain and network partition scenarios

Split-brain in a service mesh usually means both clusters believe they are authoritative and traffic is split unpredictably rather than cleanly failing over. This most often happens when the remote secret or east-west gateway is partially healthy.

Check remote secret validity

A stale or expired remote secret causes cluster-east's istiod to silently stop syncing cluster-west endpoints while reporting healthy status.

kubectl --context=cluster-east get secret -n istio-system -l istio/multiCluster=true
kubectl --context=cluster-east logs -n istio-system deploy/istiod --tail=100 | grep -i remote

Inspect endpoint consistency across clusters

Compare what each istiod instance believes about the mesh. Divergence here indicates a partition rather than a clean failover.

istioctl --context=cluster-east proxy-status
istioctl --context=cluster-west proxy-status

For deep network-level diagnosis during a partition, correlate Istio's view with raw packet flow using

Automated install script

Run this to automate the entire setup

Prefere não gerir isto sozinho?

Gerimos a infraestrutura de empresas que dependem do tempo de atividade. Totalmente gerida, com um contacto fixo que conhece o seu ambiente.

Tem um contacto fixo que conhece o seu ambiente

Na secretária em Roterdão 09:38 · acessível por mensagem, sem formulário de tickets