SSL certificate automation with Ansible lets a sysadmin generate keys, request certificates, deploy them, and reload services across hundreds of hosts from a single playbook. This guide is for teams that already use Ansible for configuration management and want certificate renewal to be one more repeatable task. It covers the modules that matter, a working renewal flow, the mistakes that cause outages, and why automation still needs an independent check from outside.
Why Certificate Renewal Belongs in Configuration Management
Certificate lifetimes are getting shorter. Under CA/Browser Forum ballot SC-081, maximum public TLS certificate validity dropped to 200 days in March 2026, falls to 100 days in March 2027, and reaches 47 days in March 2029. A team renewing 60 certificates once a year by hand will be renewing them roughly eight times a year by 2029.
Ansible fits this well because it already knows your inventory, your web server layout, and how to reload services safely. The certificate becomes declared state, just like a package version or a firewall rule. If someone deletes a file or a host gets rebuilt, the next run puts it back.
The Ansible Modules That Do the Real Work
Almost everything comes from the community.crypto collection. The openssl_privatekey module generates ECDSA or RSA keys. The openssl_csr module builds the signing request with the right Subject Alternative Names. The x509_certificate module can self-sign, sign with an internal CA through the ownca provider, or wrap an existing certificate.
For public certificates, acme_certificate talks to Let’s Encrypt, ZeroSSL, or any other ACME-compliant CA. It works in two passes: the first call returns challenge data, your playbook publishes the HTTP-01 file or DNS-01 TXT record, and the second call finishes the order. For DNS-01, pair it with community.general.cloudflare_dns, amazon.aws.route53, or your provider’s module. If you’re new to the protocol itself, the background on automating deployment with ACME is worth reading first.
Two read-only modules are used constantly. x509_certificate_info parses a certificate file on disk, and get_certificate connects to a live host and port and returns what the server actually presents. The difference between the two catches more bugs than anything else in this guide.
A Practical Renewal Playbook Step by Step
A seasoned infrastructure engineer builds the renewal flow so that it does nothing unless renewal is needed, and so that it can’t break a running service. A solid sequence looks like this:
1. Read the current certificate with x509_certificate_info and set a fact if fewer than 30 days remain (use valid_at with “+30d”).
2. Only when renewal is due, generate the CSR and run acme_certificate with the challenge pass, deploy the challenge, then complete the order.
3. Write the certificate and full chain to a staging path, validate with “nginx -t” or “apachectl configtest”, and only then move it into place.
4. Notify a handler that runs a reload, not a restart, so existing connections drain cleanly.
5. Finish with get_certificate against the public hostname and fail the play if the served serial number doesn’t match the new one.
Run it from AWX, Ansible Automation Platform, or a cron-driven CI job daily rather than weekly. A daily run on a certificate with 30 days of headroom gives you about 25 chances to recover from a failure. The same thinking applies if you trigger playbooks from pipelines; certificate management in CI/CD pipelines covers the gating side in more detail.
Private keys need special handling. Mark every task that touches key material with no_log: true, and store keys at rest with Ansible Vault or pull them from HashiCorp Vault through the community.hashi_vault lookup. Otherwise a verbose AWX job log can end up holding your private key.
Common Mistakes That Turn Automation Into Outages
The most frequent failure is deploying only the leaf certificate. acme_certificate can write the leaf, the chain, and the full chain to separate files. If nginx’s ssl_certificate points at the leaf only, desktop Chrome often still works because it fetches missing intermediates itself, but Android apps, curl, and Java clients fail. The playbook reports success while part of your user base sees errors. The guide to certificate chain issues explains how to spot this quickly.
The second mistake is trusting handler execution. Ansible handlers run at the end of a play. If a later, unrelated task fails on that host, the reload never happens. The new certificate sits on disk, x509_certificate_info says it’s valid for 199 days, and the server keeps serving the old one until it expires. Use force_handlers: true, or flush handlers with meta: flush_handlers right after deployment.
The third mistake is hitting CA rate limits in a loop. Let’s Encrypt allows 5 duplicate certificates per week for the exact same set of names. A playbook without a proper “renewal needed” condition, run from a broken retry loop, can use up that limit in an afternoon. Test against the Let’s Encrypt staging endpoint first.
A typical scenario: a team manages 40 nginx hosts with a dynamic inventory built from cloud tags. During a migration, two hosts get retagged and quietly drop out of the “webservers” group. The nightly playbook keeps reporting green because, from its point of view, every host it knows about is fine. The certificates on the two forgotten hosts expire 11 weeks later, on a Saturday.
The Myth That Automation Replaces Monitoring
The most common misconception is that once certificate renewal runs on a schedule, monitoring is redundant. It isn’t. Ansible only checks the hosts in its inventory, using the credentials and network paths it has. It has no view of a CDN edge serving a stale certificate, a load balancer holding its own copy, or a host that fell out of a group.
The external check has also become more important. Let’s Encrypt stopped sending expiration reminder emails in June 2025, so the CA no longer warns you as a fallback. An outside view that watches what clients actually receive, including chain completeness, HSTS headers, and Certificate Transparency entries for your domains, is the only thing that tests whether your automation did its job.
Edge Cases Where the Standard Approach Breaks
Air-gapped and on-prem environments can’t reach a public ACME endpoint. Here, x509_certificate with the ownca provider, or an internal ACME server like Smallstep step-ca, does the same job. The playbook structure stays identical; only the issuer changes.
Windows hosts running IIS need ansible.windows.win_certificate_store and a binding update rather than file copies. Appliances like F5 BIG-IP or Citrix ADC use their own collections, and certificates terminated on AWS ALB usually belong in ACM rather than on instances at all. Small teams of 2-3 people with fewer than ten certificates may find that certbot with a systemd timer on each host is simpler than a central playbook. Ansible makes sense when you have shared configuration, many hosts, or compliance requirements that demand an auditable change log.
FAQ
Should Ansible or certbot handle Let’s Encrypt renewals?
Either works. Certbot is simpler per host. Ansible is better when certificates must be distributed to several servers, stored centrally, or deployed together with configuration changes in one audited run.
How often should the renewal playbook run?
Daily, with renewal triggered at 30 days remaining for 90-day certificates, or at about one-third of the lifetime for shorter ones. Frequent no-op runs are cheap and give you many chances to retry before expiration.
Can Ansible verify the certificate a client actually sees?
Yes. community.crypto.get_certificate connects to the live endpoint and returns the served certificate. Compare its serial number or expiration date against the deployed file to confirm the reload actually happened.
Final Practical Tip
Treat the playbook’s own report as a claim, not proof. Make the last task of every renewal run verify the live endpoint, flush handlers right after deployment, and always deploy the full chain. Then keep an independent outside check on expiration and chain health, because the failures that hurt most are on hosts your automation doesn’t know exist.
