Build a production-grade Istio primary-remote multi-cluster topology with shared trust, locality-aware failover, and automated recovery testing. Covers cross-cluster service discovery, Kiali and Prometheus monitoring, and control plane backup procedures.
Prerequisites
- Two Kubernetes clusters (1.28+) with routable pod networks and no overlapping CIDRs
- kubectl configured with contexts for both clusters
- istioctl 1.23 or newer
- OpenSSL for CA certificate generation
- GPG for encrypted backup storage
What this solves
A single-cluster Istio mesh is a single point of failure. This tutorial configures a primary-remote multi-cluster topology spanning two Kubernetes clusters with shared trust, cross-cluster service discovery and locality-aware load balancing so traffic automatically fails over when a cluster or region goes down.
You will build the full trust chain, test real outage scenarios, wire up monitoring with Kiali and Prometheus, and establish backup and recovery procedures for the control plane configuration itself.
Prerequisites and multi-cluster architecture overview
Before starting, confirm each cluster meets the baseline requirements. Both clusters need direct pod-to-pod IP routing (no overlapping CIDRs), a shared root CA for mTLS trust, and API server access from your workstation via distinct kubeconfig contexts.
| Requirement | Details |
|---|---|
| Kubernetes version | 1.28 or newer on both clusters |
| Istio version | 1.22 or newer (istioctl and control plane must match) |
| Network | Non-overlapping pod/service CIDRs, routable between clusters |
| DNS | Cluster API servers resolvable from each other for cross-cluster secrets |
| CA | Shared intermediate CA or common root for mesh-wide mTLS trust |
The topology used here is primary-remote: cluster-east runs a full Istio control plane (istiod), and cluster-west runs a remote configuration that uses cluster-east's istiod for configuration while running its own data plane. This is simpler to operate than primary-primary and is sufficient for most disaster recovery requirements.
Install istioctl and verify cluster access
Install the Istio CLI and confirm both kubeconfig contexts are reachable before making any changes.
curl -L https://istio.io/downloadIstio | ISTIO_VERSION=1.23.2 sh -
sudo mv istio-1.23.2/bin/istioctl /usr/local/bin/
istioctl version --remote=falsekubectl config get-contexts
kubectl --context=cluster-east get nodes
kubectl --context=cluster-west get nodesSetting up primary-remote multi-cluster topology with shared trust domain
Generate a shared root CA
Every cluster in the mesh must trust the same root certificate authority so cross-cluster mTLS connections succeed. Generate an intermediate CA per cluster signed by one shared root.
mkdir -p ~/istio-ca/certs && cd ~/istio-ca/certs
curl -sL https://raw.githubusercontent.com/istio/istio/release-1.23/tools/certs/Makefile.selfsigned.mk -o Makefile
make -f Makefile root-camake -f Makefile cluster-east-cacerts
make -f Makefile cluster-west-cacertsInstall the CA secrets into each cluster
The istio-system namespace must exist before the cacerts secret is created, since istiod reads this secret at startup to establish the mesh trust domain.
kubectl --context=cluster-east create namespace istio-system
kubectl --context=cluster-east create secret generic cacerts -n istio-system \
--from-file=cluster-east/ca-cert.pem \
--from-file=cluster-east/ca-key.pem \
--from-file=cluster-east/root-cert.pem \
--from-file=cluster-east/cert-chain.pemkubectl --context=cluster-west create namespace istio-system
kubectl --context=cluster-west create secret generic cacerts -n istio-system \
--from-file=cluster-west/ca-cert.pem \
--from-file=cluster-west/ca-key.pem \
--from-file=cluster-west/root-cert.pem \
--from-file=cluster-west/cert-chain.pemLabel clusters with network and topology metadata
Istio uses the topology.istio.io/network label to determine which endpoints require the east-west gateway versus direct pod routing.
kubectl --context=cluster-east label namespace istio-system topology.istio.io/network=network-east
kubectl --context=cluster-west label namespace istio-system topology.istio.io/network=network-westInstall the primary cluster control plane
Cluster-east becomes the primary with a full istiod deployment, tagged with its mesh, network and cluster identifiers.
apiVersion: install.istio.io/v1alpha1
kind: IstioOperator
spec:
values:
global:
meshID: mesh1
multiCluster:
clusterName: cluster-east
network: network-eastistioctl install --context=cluster-east -f /tmp/cluster-east.yaml -yDeploy the east-west gateway on the primary
The east-west gateway exposes istiod and cross-cluster service endpoints over mTLS so cluster-west can reach services in cluster-east.
cd ~/istio-1.23.2
samples/multicluster/gen-eastwest-gateway.sh --mesh mesh1 --cluster cluster-east --network network-east | \
istioctl --context=cluster-east install -y -f -kubectl --context=cluster-east apply -n istio-system -f samples/multicluster/expose-istiod.yaml
kubectl --context=cluster-east apply -n istio-system -f samples/multicluster/expose-services.yamlInstall the remote cluster configuration
Cluster-west gets the east-west gateway plus a remote config that points it at cluster-east's exposed istiod service.
samples/multicluster/gen-eastwest-gateway.sh --mesh mesh1 --cluster cluster-west --network network-west | \
istioctl --context=cluster-west install -y -f -apiVersion: install.istio.io/v1alpha1
kind: IstioOperator
spec:
profile: remote
values:
istiodRemote:
injectionPath: /inject/cluster/cluster-west/net/network-west
global:
remotePilotAddress: 203.0.113.10
meshID: mesh1
multiCluster:
clusterName: cluster-west
network: network-westistioctl install --context=cluster-west -f /tmp/cluster-west.yaml -yInstall a cross-cluster remote secret
Cluster-east's istiod needs a kubeconfig secret for cluster-west so it can watch remote endpoints and services.
istioctl create-remote-secret --context=cluster-west --name=cluster-west | \
kubectl apply -f - --context=cluster-eastConfiguring cross-cluster service discovery and endpoint synchronization
With the remote secret installed, istiod on cluster-east aggregates endpoints from both clusters into a single service registry. Any workload with a matching Service name and namespace across clusters is automatically treated as a single logical service with endpoints in both locations.
Deploy a sample multi-cluster service
Deploy identical Deployment and Service manifests to both clusters. Istio merges the endpoints transparently.
apiVersion: v1
kind: Service
metadata:
name: helloworld
labels:
app: helloworld
spec:
ports:
- port: 5000
name: http
selector:
app: helloworld
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: helloworld-v1
spec:
replicas: 1
selector:
matchLabels:
app: helloworld
template:
metadata:
labels:
app: helloworld
version: v1
spec:
containers:
- name: helloworld
image: docker.io/istio/examples-helloworld-v1
ports:
- containerPort: 5000kubectl --context=cluster-east create namespace sample
kubectl --context=cluster-east label namespace sample istio-injection=enabled
kubectl --context=cluster-east apply -n sample -f /tmp/helloworld.yaml
kubectl --context=cluster-west create namespace sample
kubectl --context=cluster-west label namespace sample istio-injection=enabled
kubectl --context=cluster-west apply -n sample -f /tmp/helloworld.yamlConfirm endpoint synchronization
Query istiod's debug endpoint to confirm it sees endpoints from both clusters for the same service.
istioctl --context=cluster-east proxy-config endpoint deploy/helloworld-v1.sample -n sample | grep 5000Implementing locality-aware load balancing and failover policies
Istio's locality load balancing prefers endpoints in the same region and zone as the client, falling back to remote clusters only when local endpoints become unhealthy. This requires outlier detection to be enabled and locality labels to be set correctly on nodes.
Verify node locality labels
Istio reads region and zone from standard Kubernetes topology labels on each node.
kubectl --context=cluster-east get nodes -L topology.kubernetes.io/region,topology.kubernetes.io/zone
kubectl --context=cluster-west get nodes -L topology.kubernetes.io/region,topology.kubernetes.io/zoneCreate a DestinationRule with locality failover
This policy sends traffic to the local cluster's region first and automatically fails over to the remote region if local endpoints are ejected by outlier detection.
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
name: helloworld
namespace: sample
spec:
host: helloworld.sample.svc.cluster.local
trafficPolicy:
connectionPool:
tcp:
maxConnections: 100
outlierDetection:
consecutive5xxErrors: 3
interval: 10s
baseEjectionTime: 30s
maxEjectionPercent: 100
loadBalancer:
localityLbSetting:
enabled: true
failover:
- from: eu-west-1
to: eu-central-1kubectl --context=cluster-east apply -f /tmp/helloworld-dr.yaml
kubectl --context=cluster-west apply -f /tmp/helloworld-dr.yamlSet a failover priority with ServiceEntry (optional for external services)
If part of your mesh talks to external endpoints, define locality priority there too so DR behavior is consistent mesh-wide.
kubectl --context=cluster-east get destinationrule helloworld -n sample -o yamlThis pattern complements traffic-splitting strategies covered in Istio multi-cluster canary deployments and pairs well with the resilience patterns in Istio circuit breaker and retry policies.
Testing automated failover scenarios and cluster outage simulation
Generate baseline traffic
Use a simple loop from a client pod to confirm traffic normally stays local to cluster-east.
kubectl --context=cluster-east run sleep --image=curlimages/curl -n sample -- sleep infinity
kubectl --context=cluster-east exec -n sample sleep -- sh -c "for i in $(seq 1 20); do curl -s helloworld.sample:5000/hello; done"Simulate a cluster-east outage
Scale the local deployment to zero to simulate cluster-east becoming unavailable, then confirm traffic fails over to cluster-west.
kubectl --context=cluster-east scale deploy helloworld-v1 -n sample --replicas=0kubectl --context=cluster-east exec -n sample sleep -- sh -c "for i in $(seq 1 20); do curl -s helloworld.sample:5000/hello; done"Responses should now come from version v1 running in cluster-west. Restore the deployment once the test is complete.
kubectl --context=cluster-east scale deploy helloworld-v1 -n sample --replicas=1Simulate a full network partition
A harder test is severing the east-west gateway connection itself rather than just the workload, which exercises the outlier detection path end to end.
kubectl --context=cluster-east scale deploy istio-eastwestgateway -n istio-system --replicas=0
kubectl --context=cluster-east exec -n sample sleep -- sh -c "for i in $(seq 1 20); do curl -s -w ' -> %{http_code}\n' helloworld.sample:5000/hello; done"
kubectl --context=cluster-east scale deploy istio-eastwestgateway -n istio-system --replicas=1Monitoring multi-cluster health with Kiali and Prometheus
Install Kiali and Prometheus addons on the primary
Kiali visualizes cross-cluster traffic flow, while Prometheus scrapes per-cluster mesh metrics for alerting.
kubectl --context=cluster-east apply -f samples/addons/prometheus.yaml
kubectl --context=cluster-east apply -f samples/addons/kiali.yaml
kubectl --context=cluster-east rollout status deployment/kiali -n istio-systemConfigure remote cluster metrics scraping
Add cluster-west as a remote Prometheus target so dashboards show both clusters in one pane.
kubectl --context=cluster-west get svc istiod -n istio-systemkubectl --context=cluster-east port-forward svc/kiali 20001:20001 -n istio-systemOpen the Kiali dashboard and confirm both cluster-east and cluster-west appear under the mesh graph with cross-cluster edges between helloworld instances.
For alerting on failover events and endpoint health degradation, extend this setup with Prometheus federation for multi-cluster monitoring and route notifications through Alertmanager webhook integrations.
Backup and recovery procedures for Istio control plane configuration
Back up IstioOperator manifests and CRDs
The control plane's desired state lives in your IstioOperator YAML and any custom Istio resources (VirtualService, DestinationRule, Gateway). Export all of it regularly.
mkdir -p ~/istio-backup/$(date +%F)
cd ~/istio-backup/$(date +%F)
kubectl --context=cluster-east get istiooperator -A -o yaml > istio-operator-east.yaml
kubectl --context=cluster-east get virtualservices,destinationrules,gateways,serviceentries -A -o yaml > istio-resources-east.yaml
kubectl --context=cluster-west get virtualservices,destinationrules,gateways,serviceentries -A -o yaml > istio-resources-west.yamlBack up CA material securely
The cacerts secret is the trust root for the entire mesh. Losing it means every workload certificate must be reissued. Export it to an encrypted archive, never in plain text.
kubectl --context=cluster-east get secret cacerts -n istio-system -o yaml > cacerts-east.yaml
gpg --symmetric --cipher-algo AES256 cacerts-east.yaml
rm cacerts-east.yamlStore the resulting cacerts-east.yaml.gpg in the same offsite location used for your other encrypted secrets, following the key rotation practices in backup encryption key rotation and management.
Automate the backup with a systemd timer
Schedule the export to run daily so configuration drift never outpaces your last known-good backup.
[Unit]
Description=Istio control plane configuration backup
[Service]
Type=oneshot
ExecStart=/usr/local/bin/istio-backup.sh[Unit]
Description=Daily Istio config backup
[Timer]
OnCalendar=daily
Persistent=true
[Install]
WantedBy=timers.targetsudo systemctl daemon-reload
sudo systemctl enable --now istio-backup.timerRestore procedure
To rebuild a destroyed cluster, reinstall cacerts first, then reapply the IstioOperator manifest, then reapply custom resources.
gpg --decrypt cacerts-east.yaml.gpg > cacerts-east.yaml
kubectl --context=cluster-east apply -f cacerts-east.yaml
istioctl install --context=cluster-east -f istio-operator-east.yaml -y
kubectl --context=cluster-east apply -f istio-resources-east.yamlTroubleshooting split-brain and network partition scenarios
Split-brain in a service mesh usually means both clusters believe they are authoritative and traffic is split unpredictably rather than cleanly failing over. This most often happens when the remote secret or east-west gateway is partially healthy.
Check remote secret validity
A stale or expired remote secret causes cluster-east's istiod to silently stop syncing cluster-west endpoints while reporting healthy status.
kubectl --context=cluster-east get secret -n istio-system -l istio/multiCluster=true
kubectl --context=cluster-east logs -n istio-system deploy/istiod --tail=100 | grep -i remoteInspect endpoint consistency across clusters
Compare what each istiod instance believes about the mesh. Divergence here indicates a partition rather than a clean failover.
istioctl --context=cluster-east proxy-status
istioctl --context=cluster-west proxy-statusFor deep network-level diagnosis during a partition, correlate Istio's view with raw packet flow using
Automated install script
Run this to automate the entire setup
#!/usr/bin/env bash
set -euo pipefail
# ---------------------------------------------------------------------------
# Istio multi-cluster (primary-remote) disaster recovery bootstrap script
# Configures shared trust, installs istioctl, and prepares CA material
# for cluster-east (primary) and cluster-west (remote).
# ---------------------------------------------------------------------------
ISTIO_VERSION="${ISTIO_VERSION:-1.23.2}"
CTX_EAST="${1:-cluster-east}"
CTX_WEST="${2:-cluster-west}"
WORKDIR="${WORKDIR:-$HOME/istio-ca}"
INSTALL_DIR="/usr/local/bin"
TOTAL_STEPS=9
RED='\033[0;31m'
GREEN='\033[0;32m'
YELLOW='\033[1;33m'
NC='\033[0m'
log() { echo -e "${GREEN}[$1/${TOTAL_STEPS}]${NC} $2"; }
warn() { echo -e "${YELLOW}WARNING:${NC} $1"; }
err() { echo -e "${RED}ERROR:${NC} $1" >&2; }
usage() {
echo "Usage: $0 [cluster-east-context] [cluster-west-context]"
echo " Defaults: cluster-east, cluster-west"
echo " Requires: kubectl configured with both contexts reachable"
exit 1
}
if [[ "${1:-}" == "-h" || "${1:-}" == "--help" ]]; then
usage
fi
CLEANUP_NEEDED=0
cleanup() {
if [ "$CLEANUP_NEEDED" -eq 1 ]; then
err "Install failed. Rolling back partially created resources..."
kubectl --context="$CTX_EAST" delete secret cacerts -n istio-system --ignore-not-found=true >/dev/null 2>&1 || true
kubectl --context="$CTX_WEST" delete secret cacerts -n istio-system --ignore-not-found=true >/dev/null 2>&1 || true
warn "Namespaces istio-system were left in place to avoid destroying unrelated workloads."
warn "Review ${WORKDIR} for generated certificate material before retrying."
fi
}
trap cleanup ERR
# ---------------------------------------------------------------------------
# [1/N] Detect distro and package manager
# ---------------------------------------------------------------------------
if [ -f /etc/os-release ]; then
. /etc/os-release
case "$ID" in
ubuntu|debian) PKG_MGR="apt"; PKG_INSTALL="apt install -y" ;;
almalinux|rocky|centos|rhel|ol|fedora) PKG_MGR="dnf"; PKG_INSTALL="dnf install -y" ;;
amzn) PKG_MGR="yum"; PKG_INSTALL="yum install -y" ;;
*) err "Unsupported distro: $ID"; exit 1 ;;
esac
else
err "/etc/os-release not found. Cannot determine distro."
exit 1
fi
log 1 "Detected distro: $ID (package manager: $PKG_MGR)"
# ---------------------------------------------------------------------------
# [2/N] Check prerequisites: root/sudo and required tools
# ---------------------------------------------------------------------------
log 2 "Checking prerequisites..."
if [ "$(id -u)" -ne 0 ]; then
if ! command -v sudo >/dev/null 2>&1; then
err "This script must be run as root or with sudo available."
exit 1
fi
SUDO="sudo"
else
SUDO=""
fi
# Install base tools if missing (curl, make, tar, kubectl assumed pre-existing)
MISSING_PKGS=()
for tool in curl make tar; do
if ! command -v "$tool" >/dev/null 2>&1; then
MISSING_PKGS+=("$tool")
fi
done
if [ "${#MISSING_PKGS[@]}" -gt 0 ]; then
echo "Installing missing dependencies: ${MISSING_PKGS[*]}"
if [ "$PKG_MGR" = "apt" ]; then
$SUDO apt update -y
fi
$SUDO $PKG_INSTALL "${MISSING_PKGS[@]}"
fi
if ! command -v kubectl >/dev/null 2>&1; then
err "kubectl is required but not installed. Install kubectl before running this script."
exit 1
fi
# ---------------------------------------------------------------------------
# [3/N] Verify both kubeconfig contexts are reachable
# ---------------------------------------------------------------------------
log 3 "Verifying kubeconfig contexts '$CTX_EAST' and '$CTX_WEST'..."
if ! kubectl config get-contexts "$CTX_EAST" >/dev/null 2>&1; then
err "Kube context '$CTX_EAST' not found in kubeconfig."
usage
fi
if ! kubectl config get-contexts "$CTX_WEST" >/dev/null 2>&1; then
err "Kube context '$CTX_WEST' not found in kubeconfig."
usage
fi
if ! kubectl --context="$CTX_EAST" get nodes >/dev/null 2>&1; then
err "Cannot reach cluster via context '$CTX_EAST'."
exit 1
fi
if ! kubectl --context="$CTX_WEST" get nodes >/dev/null 2>&1; then
err "Cannot reach cluster via context '$CTX_WEST'."
exit 1
fi
echo "Both clusters reachable."
# ---------------------------------------------------------------------------
# [4/N] Install istioctl (idempotent)
# ---------------------------------------------------------------------------
log 4 "Installing istioctl v${ISTIO_VERSION}..."
if command -v istioctl >/dev/null 2>&1 && istioctl version --remote=false 2>/dev/null | grep -q "$ISTIO_VERSION"; then
echo "istioctl ${ISTIO_VERSION} already installed, skipping."
else
TMP_ISTIO_DIR="$(mktemp -d)"
pushd "$TMP_ISTIO_DIR" >/dev/null
curl -sL "https://istio.io/downloadIstio" | ISTIO_VERSION="$ISTIO_VERSION" sh -
$SUDO install -m 755 -o root -g root "istio-${ISTIO_VERSION}/bin/istioctl" "${INSTALL_DIR}/istioctl"
popd >/dev/null
rm -rf "$TMP_ISTIO_DIR"
fi
istioctl version --remote=false
# ---------------------------------------------------------------------------
# [5/N] Generate shared root CA and per-cluster intermediate CAs
# ---------------------------------------------------------------------------
log 5 "Generating shared root CA and intermediate certs..."
mkdir -p "${WORKDIR}/certs"
chmod 755 "$WORKDIR" "${WORKDIR}/certs"
cd "${WORKDIR}/certs"
if [ ! -f Makefile ]; then
curl -sL https://raw.githubusercontent.com/istio/istio/release-1.23/tools/certs/Makefile.selfsigned.mk -o Makefile
chmod 644 Makefile
fi
if [ ! -f root-cert.pem ]; then
make -f Makefile root-ca
fi
if [ ! -d "${CTX_EAST}" ]; then
make -f Makefile "${CTX_EAST}-cacerts"
fi
if [ ! -d "${CTX_WEST}" ]; then
make -f Makefile "${CTX_WEST}-cacerts"
fi
# Restrict private key material to owner only
find "${WORKDIR}/certs" -name '*-key.pem' -exec chmod 600 {} \;
# ---------------------------------------------------------------------------
# [6/N] Create istio-system namespaces and install cacerts secrets
# ---------------------------------------------------------------------------
log 6 "Creating namespaces and installing CA secrets..."
CLEANUP_NEEDED=1
for ctx in "$CTX_EAST" "$CTX_WEST"; do
kubectl --context="$ctx" get namespace istio-system >/dev/null 2>&1 || \
kubectl --context="$ctx" create namespace istio-system
if kubectl --context="$ctx" get secret cacerts -n istio-system >/dev/null 2>&1; then
warn "cacerts secret already exists in context '$ctx', skipping creation."
else
kubectl --context="$ctx" create secret generic cacerts -n istio-system \
--from-file="${ctx}/ca-cert.pem" \
--from-file="${ctx}/ca-key.pem" \
--from-file="${ctx}/root-cert.pem" \
--from-file="${ctx}/cert-chain.pem"
fi
done
# ---------------------------------------------------------------------------
# [7/N] Label namespaces with network topology metadata
# ---------------------------------------------------------------------------
log 7 "Labeling clusters with topology.istio.io/network..."
kubectl --context="$CTX_EAST" label namespace istio-system \
topology.istio.io/network=network-east --overwrite
kubectl --context="$CTX_WEST" label namespace istio-system \
topology.istio.io/network=network-west --overwrite
# ---------------------------------------------------------------------------
# [8/N] Configure firewall for east-west gateway traffic (if active)
# ---------------------------------------------------------------------------
log 8 "Checking local firewall configuration for Istio gateway ports..."
ISTIO_PORTS=(15021 15443 15012 15017)
if command -v firewall-cmd >/dev/null 2>&1 && $SUDO systemctl is-active --quiet firewalld 2>/dev/null; then
for port in "${ISTIO_PORTS[@]}"; do
$SUDO firewall-cmd --permanent --add-port="${port}/tcp" >/dev/null
done
$SUDO firewall-cmd --reload
echo "firewalld updated for ports: ${ISTIO_PORTS[*]}"
elif command -v ufw >/dev/null 2>&1 && $SUDO ufw status 2>/dev/null | grep -q "Status: active"; then
for port in "${ISTIO_PORTS[@]}"; do
$SUDO ufw allow "${port}/tcp" >/dev/null
done
echo "ufw updated for ports: ${ISTIO_PORTS[*]}"
else
warn "No active local firewall detected (firewalld/ufw). Ensure cloud security groups allow ports: ${ISTIO_PORTS[*]}"
fi
# ---------------------------------------------------------------------------
# [9/N] Verification
# ---------------------------------------------------------------------------
log 9 "Running verification checks..."
FAILED=0
for ctx in "$CTX_EAST" "$CTX_WEST"; do
if kubectl --context="$ctx" get secret cacerts -n istio-system >/dev/null 2>&1; then
echo " [OK] cacerts present in $ctx"
else
err " [FAIL]
Review the script before running. Execute with: bash install.sh