+1 909 277 6076 support@otsglobal.org

A real-world troubleshooting session on a Rocky Linux 10.2 iRedMail server running under KVM


The Setup

A small mail server hosted on a KVM virtualization host. The host runs Rocky Linux 10.2 and manages two bridges: one for the public-facing network and one for the internal LAN. The mail server itself is a VM running iRedMail on Rocky Linux 10.2, with Dovecot, Postfix, and Nginx serving IMAP, SMTP, and webmail respectively.

DNS is split-horizon: an internal Windows DNS server returns the LAN IP for the mail hostname, while public DNS returns the public IP. This lets LAN clients talk to the mail server directly, without hairpin NAT.

Everything was working. Until it wasn't.

The Symptoms

Users started reporting that Thunderbird could not connect. The error was classic and unhelpful:

"Could not connect to server; the connection was refused."

External clients (from the internet) could still connect. LAN clients could not. That asymmetry is always a clue.

At the same time, a recent change had been made to Dovecot to tighten security — disabling plaintext authentication to force clients onto TLS. That change was the trigger, but not the root cause.

The Investigation

We approached this systematically, one layer at a time.

Layer 1: Is anything listening?

The first check was simple:

ss -tlnp | grep -E ':(110|143|993|995)'

Dovecot was listening on 0.0.0.0 on all four mail ports. That ruled out a binding issue — the service was reachable on every interface.

Layer 2: Does the LAN path work?

From the KVM host, we tested connectivity to the mail VM's LAN IP:

ping -c 3 <lan-ip>
nc -zv <lan-ip> 993
nc -zv <lan-ip> 143
nc -zv <lan-ip> 25

All three ports responded. The bridge, the firewall zone, and the mail VM's own firewall were all fine. That ruled out the network path.

Layer 3: What does the LAN client resolve?

From a LAN client, DNS returned the public IP, not the LAN IP. That would normally explain the failure — LAN clients trying to reach the public IP require hairpin NAT, which most routers don't do.

But the internal Windows DNS server did return the LAN IP when queried directly:

nslookup mail.example.com <internal-dns-ip>

That worked. So the problem was that LAN clients weren't using the internal DNS. That was one issue, but not the whole story.

Layer 4: Does TLS actually work?

From the mail server itself:

openssl s_client -connect 127.0.0.1:993 -servername mail.example.com

The handshake failed. No certificate was being served. This was the real problem — TLS was broken server-side, so no client could connect, regardless of DNS.

The Root Cause

A quick check of the certificate key permissions revealed everything:

ls -la /etc/letsencrypt/archive/mail.example.com/privkey*.pem
-rw------- 1 root root 241 ... privkey1.pem
-rw------- 1 root root 241 ... privkey2.pem
...

The private key was root:root 0600. Dovecot, Postfix, and Nginx all run as unprivileged users after startup. None of them could read the key. They started successfully but silently refused all TLS connections.

The trigger: a recent Let's Encrypt renewal had reset the key permissions to root-only. There was no renewal hook to fix them back. And the recent Dovecot ssl = required change meant clients could no longer fall back to plaintext — they needed TLS, and TLS was unavailable.

The combination of (a) an existing latent bug (unreadable key after renewal) and (b) a recent security tightening (requiring TLS) produced the outage.

The Fix

We created a dedicated group for certificate readers, added the three services to it, and granted group read access to the private key:

groupadd -f sslcerts
usermod -aG sslcerts dovecot
usermod -aG sslcerts postfix
usermod -aG sslcerts nginx

chmod 755 /etc/letsencrypt/archive/mail.example.com
chgrp sslcerts /etc/letsencrypt/archive/mail.example.com/privkey*.pem
chmod 640 /etc/letsencrypt/archive/mail.example.com/privkey*.pem

systemctl restart dovecot postfix nginx

Verified:

sudo -u dovecot cat /etc/letsencrypt/live/mail.example.com/privkey.pem > /dev/null && echo OK
sudo -u postfix cat /etc/letsencrypt/live/mail.example.com/privkey.pem > /dev/null && echo OK
sudo -u nginx   cat /etc/letsencrypt/live/mail.example.com/privkey.pem > /dev/null && echo OK

All three reported OK.

Then a real TLS handshake test:

openssl s_client -connect 127.0.0.1:993 -servername mail.example.com

The certificate chain appeared. Verify return code: 0 (ok).

From an external machine, we tested all five protocols:

openssl s_client -connect mail.example.com:993 -servername mail.example.com    # IMAPS
openssl s_client -connect mail.example.com:465 -servername mail.example.com    # SMTPS
openssl s_client -connect mail.example.com:143 -starttls imap -servername mail.example.com
openssl s_client -connect mail.example.com:587 -starttls smtp -servername mail.example.com
curl -sI https://mail.example.com

All returned valid certificates. TLS was restored.

Preventing Recurrence

The critical part: fixing permissions once is not enough. Let's Encrypt renews certificates every 60–90 days, and certbot resets the key permissions to root:root 0600 on every renewal. Without a hook to fix them, the same outage recurs on the next cycle.

We enhanced the existing renewal hook:

#!/bin/bash
# /etc/letsencrypt/renewal-hooks/post/renew-mail.sh

ARCHIVE_DIR="/etc/letsencrypt/archive/mail.example.com"

# Certbot resets key permissions on renewal — fix them before reloading services
if [ -d "$ARCHIVE_DIR" ]; then
    chmod 755 "$ARCHIVE_DIR"
    chgrp sslcerts "$ARCHIVE_DIR"/privkey*.pem 2>/dev/null || true
    chmod 640 "$ARCHIVE_DIR"/privkey*.pem 2>/dev/null || true
fi

systemctl reload nginx
systemctl reload postfix
systemctl reload dovecot

echo "$(date): Certificate renewed, permissions fixed, services reloaded" >> /var/log/letsencrypt-renewal.log

Made executable, tested directly, and confirmed with a dry-run:

certbot renew --dry-run --cert-name mail.example.com

The hook fires on every renewal, fixes permissions, and reloads services automatically.

The Split-Horizon DNS Fix

Separately, we addressed the LAN DNS issue. The mail server's own firewall on the KVM host was correct, but the host's LAN bridge was in the wrong zone — public instead of trusted — which is a common oversight when a KVM host serves both public and internal networks.

The fix:

nmcli connection modify br0 connection.zone trusted
nmcli connection modify br1 connection.zone public
systemctl restart firewalld

And, importantly, we cleaned up the stale <interface> reference from the zone XML files. On RHEL 10 / firewalld 2.x, NetworkManager owns interface-to-zone assignment, and having that information duplicated in XML can cause the daemon to fail on reload with cryptic errors like dictionary changed size during iteration.

After the fix:

public
  interfaces: br1 eno1 eno2
trusted
  interfaces: br0

Internal traffic is trusted and unrestricted; external traffic stays under public rules.

Lessons Learned

1. TLS failures often look like connection failures.

"Connection refused" from a mail client is not always a network or firewall problem. When a service has ssl = required, and the certificate can't be loaded, the service accepts the TCP connection but immediately closes it at the TLS handshake. Clients report this as "refused."

2. Certificate permissions are a recurring problem, not a one-time fix.

Certbot resets permissions on renewal. If your services don't run as root, you need a renewal hook — always. The default renewal-hooks/post directory is the right place.

3. Use a dedicated group for cert readers.

Adding services to dovecot's group is lazy. A dedicated sslcerts group makes the intent explicit, avoids granting unrelated privileges, and is easy to audit.

4. Test with openssl s_client, not just service status.

A service can report active (running) while being completely unable to serve TLS. The only real test is a handshake:

openssl s_client -connect 127.0.0.1:993 -servername mail.example.com

5. Split-horizon DNS requires client-side cooperation.

Even with a perfect internal DNS server, clients that don't use it will resolve external IPs and fail. Verify each client's actual DNS resolver, and fix DHCP if needed.

6. Firewalld on RHEL 10 is NM-driven.

Interface-to-zone assignments live in NetworkManager, not in firewalld's XML. Modifying zones via firewall-cmd while NM is also managing them can produce race conditions. Use nmcli connection modify and keep the zone XML files clean.

7. Save a diagnostic script.

The whole investigation took an hour. A simple check-tls.sh would have shown the problem in seconds:

#!/bin/bash
systemctl is-active dovecot postfix nginx
ls -la /etc/letsencrypt/archive/mail.example.com/privkey*.pem
openssl x509 -in /etc/letsencrypt/live/mail.example.com/fullchain.pem -noout -dates
openssl s_client -connect 127.0.0.1:993 -servername mail.example.com < /dev/null 2>&1 | grep 'Verify return'

Any time a mail client complains, run it first.

Post-Mortem Summary

LayerFindingFix
DNSLAN clients resolved to public IPConfirmed split-horizon; fix DHCP to serve internal DNS
Firewall (KVM host)LAN bridge in wrong zoneMoved br0 to trusted via NetworkManager
Firewall (mail VM)OK — ports openNo change
Service bindingOK — listening on all interfacesNo change
TLSPrivate key unreadable by servicessslcerts group + permission fix
RenewalNo hook to fix permissions after renewalEnhanced renew-mail.sh hook

The outage was caused by a latent misconfiguration (key permissions) triggered by a legitimate security change (requiring TLS). Both are now fixed permanently, with the renewal hook closing the loop for future cycles.


If you run a mail server on RHEL-family Linux with Let's Encrypt, do yourself a favor: check your private key permissions today. If they're root:root 0600, and your mail services don't run as root, you're one renewal away from a very confusing outage.