What you will achieve and why it matters
By the end of this guide you will have a fleet of servers that renew their TLS certificates automatically, reload the correct services without manual intervention, and alert you only when something actually needs attention. This matters beyond uptime. Manual certificate tracking across many hosts is a recurring labor cost, and it is exactly the kind of recurring manual task that shows up in audits of cloud cost optimization services engagements: someone on the team spends a few hours a month on something that should take zero.
We will build this using Let's Encrypt as the certificate authority, the ACME protocol for automation, and a central orchestration pattern that works whether you are running 5 servers or 500. The same pattern applies whether your fleet sits behind a load balancer, runs as separate application nodes, or spans multiple regions as part of a broader high availability infrastructure setup.
Prerequisites and assumptions
This guide assumes:
- You control DNS for the domains you are issuing certificates for, or at minimum can create TXT records programmatically via an API (Cloudflare, Route 53, DigitalOcean, etc.)
- Your servers run Linux (examples use Ubuntu 22.04/24.04, but the approach is distro-agnostic)
- You have SSH access to all fleet members, ideally via a config management tool (Ansible, Salt, or similar) rather than manual SSH loops
- Services that terminate TLS are Nginx, HAProxy, or a similar reverse proxy, not an application server directly handling certificates
- You are comfortable with cron or systemd timers for scheduling
We will use certbot as the ACME client because it has the widest plugin support and is well maintained, but the same architecture works with acme.sh or lego if you prefer a lighter footprint.
Step-by-step implementation
Step 1: Choose a validation method that scales
HTTP-01 validation requires port 80 to be reachable on each host for the specific domain being validated. This works for single-domain setups but becomes painful across a fleet, especially behind a load balancer where traffic for a given domain might not land on a predictable server.
DNS-01 validation is the better default for fleets. It does not require inbound traffic to reach the renewing server, supports wildcard certificates, and lets you centralize certificate issuance on one host even if the certificate is then distributed to many.
Example using Cloudflare as the DNS provider:
apt install certbot python3-certbot-dns-cloudflare -y
cat > /etc/letsencrypt/cloudflare.ini << 'EOF'
dns_cloudflare_api_token = YOUR_SCOPED_API_TOKEN
EOF
chmod 600 /etc/letsencrypt/cloudflare.iniUse a scoped API token (Zone:DNS:Edit for the specific zone only), not your global API key. If that token leaks, the blast radius is limited to DNS records, not your entire Cloudflare account.
Step 2: Issue the certificate from a single orchestration host
Rather than running certbot independently on every server (which risks rate limit collisions and inconsistent certificate state), designate one host, or a small automation runner, as the certificate authority client.
certbot certonly \
--dns-cloudflare \
--dns-cloudflare-credentials /etc/letsencrypt/cloudflare.ini \
--dns-cloudflare-propagation-seconds 30 \
-d example.com \
-d '*.example.com' \
--agree-tos \
-m ops@example.com \
--non-interactiveThis produces a wildcard certificate covering all subdomains, which is usually the right call for a fleet serving multiple services under one domain. If your services span unrelated domains, issue separate certificates per domain rather than cramming everything into one SAN certificate. It keeps the blast radius of a renewal failure smaller.
Step 3: Distribute certificates to fleet members
Once issued, certificates live in /etc/letsencrypt/live/example.com/ on the orchestration host. You need a reliable way to push the fullchain.pem and privkey.pem to every server that terminates TLS.
An Ansible playbook handles this cleanly:
- name: Distribute TLS certificates to fleet
hosts: web_fleet
tasks:
- name: Copy fullchain
copy:
src: /etc/letsencrypt/live/example.com/fullchain.pem
dest: /etc/ssl/certs/example.com.pem
owner: root
group: root
mode: '0644'
notify: reload nginx
- name: Copy private key
copy:
src: /etc/letsencrypt/live/example.com/privkey.pem
dest: /etc/ssl/private/example.com.key
owner: root
group: root
mode: '0600'
notify: reload nginx
handlers:
- name: reload nginx
systemd:
name: nginx
state: reloadedThe handler only fires if the file content actually changed, which means idle runs do not cause unnecessary reloads. Nginx reload is graceful, so in-flight connections are not dropped.
Step 4: Automate renewal and redistribution on a schedule
Certbot ships with a systemd timer that runs twice daily and only renews certificates within 30 days of expiry, which is the right cadence. But the default renewal hook only runs locally. You need a deploy hook that triggers your distribution playbook.
mkdir -p /etc/letsencrypt/renewal-hooks/deploy
cat > /etc/letsencrypt/renewal-hooks/deploy/distribute.sh << 'EOF'
#!/bin/bash
set -euo pipefail
ansible-playbook -i /etc/ansible/hosts /opt/playbooks/distribute-certs.yml
EOF
chmod +x /etc/letsencrypt/renewal-hooks/deploy/distribute.shDeploy hooks only run when a certificate actually renews, unlike pre/post hooks which run on every invocation. This is the correct hook type for triggering fleet-wide distribution, since you do not want to run an Ansible playbook twice a day for no reason.
Verify the timer is active:
systemctl status certbot.timer
systemctl list-timers | grep certbotStep 5: Handle non-Nginx services
If some fleet members run HAProxy, which expects a combined PEM file (cert plus key concatenated), add a step to the deploy hook:
cat /etc/letsencrypt/live/example.com/fullchain.pem \
/etc/letsencrypt/live/example.com/privkey.pem \
> /etc/haproxy/certs/example.com.pem
systemctl reload haproxyKeep per-service post-processing logic in the deploy hook script rather than scattering it across playbooks. One script, one source of truth for what happens after a renewal.
Step 6: Centralize logging and failure alerting
A renewal that fails silently is worse than no automation at all, because it creates false confidence. Pipe certbot's renewal logs to a place you actually look:
certbot renew --dry-run --quiet 2>&1 | logger -t certbot-renewalFor actual alerting, wrap the renewal command and check the exit code:
#!/bin/bash
if ! certbot renew --quiet; then
curl -X POST https://hooks.slack.com/services/YOUR/WEBHOOK/URL \
-d '{"text":"Certbot renewal failed on '"$(hostname)"'"}'
fiReplace the webhook call with whatever your team already uses for alerting (PagerDuty, Opsgenie, a Slack channel), the important part is that a failed renewal generates a signal somewhere a human will see it within a day, not 29 days later when the certificate is about to expire.
Verification: how to confirm it works
Do not assume automation works because the cron job exists. Verify it end to end.
Dry-run the renewal process:
certbot renew --dry-runThis simulates the full renewal flow, including DNS validation, without hitting Let's Encrypt's production rate limits. Run it immediately after setup and again after any configuration change.
Check certificate expiry dates across the fleet:
for host in $(cat fleet_hosts.txt); do
echo -n "$host: "
echo | openssl s_client -connect $host:443 -servername $host 2>/dev/null \
| openssl x509 -noout -enddate
doneRun this weekly, or better, wire it into your existing monitoring stack as a scheduled check. Most monitoring tools (Prometheus with the blackbox exporter, Uptime Kuma, Zabbix) have a built-in certificate expiry probe.
Prometheus blackbox exporter example:
- job_name: 'ssl_expiry'
metrics_path: /probe
params:
module: [http_2xx]
static_configs:
- targets:
- https://example.com
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- target_label: __address__
replacement: blackbox-exporter:9115Alert on probe_ssl_earliest_cert_expiry dropping below a 14-day threshold. This gives you a second, independent confirmation layer outside of certbot's own logging, which matters because you want to know if the automation itself breaks, not just whether the last run succeeded.
Confirm distribution actually reached every host:
ansible web_fleet -m shell -a "openssl x509 -in /etc/ssl/certs/example.com.pem -noout -enddate"Run this across the fleet after a renewal cycle and confirm every host reports the same expiry date. A mismatch means your distribution step failed on at least one node, which is a far more common failure mode than the renewal itself failing.
Common pitfalls to avoid
- Hardcoding API tokens in playbooks. Use a secrets manager (Ansible Vault, HashiCorp Vault, or your cloud provider's secret store) instead of committing credentials to a repo.
- Forgetting to reload services after distribution. Copying a new certificate file does nothing if Nginx or HAProxy is still holding the old one in memory. Always pair file updates with a reload handler.
- Hitting Let's Encrypt rate limits. The production rate limit is 50 certificates per registered domain per week. If you are issuing per-subdomain certificates across a large fleet instead of one wildcard, you can hit this faster than expected. Test against the staging environment (
--stagingflag) while debugging. - Running certbot independently on every server. This creates inconsistent certificate state and multiplies the number of things that can silently break. Centralize issuance, distribute the result.
- No alerting on renewal failure. A cron job that fails quietly for a month is not automation, it is a delayed incident.
Next steps and related reading
Once certificate renewal is automated, the same orchestration pattern, central issuance with fleet-wide distribution via config management, applies to other recurring operational tasks: rotating SSH keys, pushing updated firewall rules, or redeploying configuration after a zero-downtime migration.
If your fleet sits behind a load balancer with multiple edge nodes, it is worth reviewing how your server performance holds under real traffic once TLS termination is handled consistently across every node, since certificate handling and connection handling often share the same Nginx or HAProxy configuration files.
Teams running this at scale often fold certificate automation into a broader GitOps workflow, where certificate distribution is just one more state reconciled by the same pipeline that manages application deployments. That reduces the number of distinct automation systems you have to maintain and monitor separately.
Close
Automating certificate renewal removes a recurring manual task from your team's plate and closes a gap that otherwise shows up as unplanned downtime or last-minute scrambling. It is a small piece of infrastructure, but it is exactly the kind of small piece that compounds across a fleet.
Need this running in production without building it yourself? See our managed infrastructure services or schedule a call.