Last week I got a message that the CEO at one of our client sites couldn't send or receive email when he was in the office. Anywhere else, everything worked. Same laptop, same account, same server. Only on the office LAN did mail just... not go anywhere.
I've seen this shape of problem before and it's usually DNS. Sometimes it's a firewall. Occasionally it's something weird with how the client authenticates on a particular network. This one ended up being two of the three, and chasing the second one led me to rip out Fail2ban entirely, which I'd been meaning to do for a while anyway.
First thing: what does his Mac actually think the mail server is
Ran a quick dig from his machine while he was in the office.
$ dig +short mail.example.com
203.0.113.45
The mail server lives at 203.0.113.100. The office DNS was handing out a completely different address. Not even close. That address hadn't hosted the mail server in years but it was still sitting in the office DNS as an override someone had set up who knows when.
So his Mac was connecting to a dead host. Nothing on the mail server to see, because the mail server never got contacted. Easy fix, updated the office DNS to point at the right place, or actually just removed the override and let public DNS do its thing.
Thought we were done. He tried again from the office. Still failed.
Second thing: the ports that work and the ports that don't
Had him run a port test from his Mac.
$ for p in 465 587 993; do nc -vz -G 3 mail.example.com $p 2>&1 | tail -1; done
nc: connectx to mail.example.com port 465 failed: Connection refused
nc: connectx to mail.example.com port 587 failed: Connection refused
Connection to mail.example.com port 993 succeeded!
IMAP works. SMTP submission doesn't. That's a really specific pattern. If the office router was blocking all outbound mail it would block IMAP too. So something on our end is filtering specific ports.
My first guess was the office firewall blocking SMTP as policy, which is common enough. But I wanted to check our side first before blaming the client's network.
I already knew from the handoff notes that this server had been running Fail2ban for years and had a CrowdSec install sitting on top of it. Both were active at the same time. That overlap should have raised a red flag earlier than it did.
Checked Fail2ban's internal state for the office IP.
$ OFFICE_WAN=198.51.100.10
$ fail2ban-client banned | tr ',' '\n' | grep -F "$OFFICE_WAN"
not banned in any jail
Not banned. So it's not Fail2ban, right?
Except fail2ban-client banned only shows what Fail2ban thinks is banned. It doesn't show what's actually in the kernel. That distinction is the whole ballgame here.
Looked at the actual iptables rules.
$ iptables -L -n -v | grep -E '465|587'
f2b-postfix-burst tcp -- * * 0.0.0.0/0 0.0.0.0/0 multiport dports 25,587,465
f2b-postfix-spam tcp -- * * 0.0.0.0/0 0.0.0.0/0 multiport dports 25,587,465
f2b-postfix tcp -- * * 0.0.0.0/0 0.0.0.0/0 multiport dports 25,587,465
Six Fail2ban chains, each holding hundreds of individual reject rules. About 7,000 rules total. And in there, rule 4 of the postfix chain:
REJECT all -- * * 198.51.100.10 0.0.0.0/0 reject-with icmp-port-unreachable
The office IP was banned in the kernel but not in Fail2ban's own state. This is a thing that happens on Rocky and other modern distros where the kernel has both iptables and nftables and Fail2ban is using the iptables-compat layer while firewalld is running on nftables. Fail2ban adds the rule, the ban expires or gets cleared from its database, but the actual rule in the kernel never gets removed. It just sits there forever.
And because it's a reject on the kernel level, the packet never reaches Postfix. Postfix has no idea any of this is happening. Nothing in the mail logs. Nothing in the Fail2ban logs. Just a silent-ish rejection at the packet level.
The port pattern makes perfect sense now too. Ports 465 and 587 are in f2b-postfix, which had banned the office IP. Port 993 is only in f2b-dovecot, which had never banned it, so IMAP was fine. Everything lines up.
The fix, and the realization that this would keep happening
Unbanned the affected IPs from every jail, then flushed the chains manually because Fail2ban's unban doesn't actually touch the kernel rules on this setup.
for ip in 198.51.100.10 203.0.113.100; do
for j in $(fail2ban-client status | awk -F: '/Jail list/{print $2}' | tr ',' ' '); do
fail2ban-client set "$j" unbanip "$ip" 2>/dev/null
done
done
for c in f2b-postfix f2b-postfix-burst f2b-postfix-spam f2b-dovecot f2b-roundcube f2b-pregreet; do
iptables -F "$c"
iptables -X "$c"
done
CEO's mail worked immediately after that. Ticket closed, technically.
But I sat there looking at the chain of events and I knew we'd be back here. This bug isn't a one-off. It's structural. Fail2ban writing to iptables while firewalld owns nftables, two ban systems running side by side with separate whitelists that don't know about each other, and a non-atomic unban that leaves garbage in the kernel. The next time any trusted IP tripped a jail, we'd get another orphaned rule and another mystery outage.
We'd already started moving to CrowdSec a while back. It had been running alongside Fail2ban as a safety net. Time to actually finish the migration.
Before pulling the trigger, one sanity check
CrowdSec was enforcing some bans already, Fail2ban was enforcing a lot more. If I turned off Fail2ban and CrowdSec wasn't covering the same traffic, I'd have a real gap. Quick gap analysis.
fail2ban-client banned | tr -d "[]'" | tr ',' '\n' | \
grep -v '^ ' | grep -v '^$' | sort -u > /tmp/f2b-ips.txt
cscli decisions list --origin crowdsec -o json | \
grep -oE '"value": *"[^"]+"' | cut -d'"' -f4 | sed 's/^Ip://' | sort -u > /tmp/cs-ips.txt
comm -23 /tmp/f2b-ips.txt /tmp/cs-ips.txt
One IP in Fail2ban that CrowdSec didn't have. A single SSH attempt that hadn't crossed CrowdSec's threshold. Effectively full coverage. Good enough to proceed.
The cutover
Snapshotted everything first, then stopped and masked Fail2ban, removed its cron entry, moved the DB script to a retirement folder, flushed and deleted every f2b chain, and removed the jump rules from INPUT.
tar czf /root/snapshots/pre-f2b-retire-$(date +%F-%H%M).tar.gz \
/etc/fail2ban/ /etc/crowdsec/ /etc/firewalld/ /usr/local/bin/ /var/spool/cron/root
systemctl stop fail2ban
systemctl disable fail2ban
systemctl mask fail2ban
crontab -l | grep -v 'fail2ban_banned_db' | crontab -
mv /usr/local/bin/fail2ban_banned_db /root/retired-scripts/immediate-$(date +%F)/
for c in $(iptables -L -n | grep -oE 'f2b-[a-z-]+' | sort -u); do
iptables -F "$c" 2>/dev/null
iptables -X "$c" 2>/dev/null
done
while iptables -L INPUT -n --line-numbers | grep -q 'f2b-'; do
line=$(iptables -L INPUT -n --line-numbers | grep 'f2b-' | head -1 | awk '{print $1}')
iptables -D INPUT "$line"
done
Verified after: Fail2ban masked, inactive, no processes. iptables -L | grep f2b returns nothing. Cron clean. CrowdSec still running with its decisions intact. Mail services up. All good.
One improvement while I was in there
CrowdSec has a feature that's the modern equivalent of Fail2ban's recidive jail — escalating ban times for repeat offenders. It's a single commented line in profiles.yaml.
duration_expr: Sprintf('%dh', (GetDecisionsCount(Alert.GetValue()) + 1) * 4)
First offense gets 4 hours, second gets 8, third 12, and so on. We had a Romanian SSH netblock that got 4-hour bans the day before. They came back the next day and got 144-hour bans. Six days. Exactly what you want.
The whitelist thing, which almost bit us
One thing Fail2ban did that I needed CrowdSec to do properly was the whitelist. In Fail2ban it's ignoreip. In CrowdSec it's a whitelist parser in s02-enrich. You list the IPs and ranges that should never be touched — office WAN, admin IPs, server's own subnet, that sort of thing.
I added the entries, reloaded, and it looked fine. But I got suspicious. "Looks fine" and "is fine" are different things when it comes to whitelists. If the whitelist is loaded but broken, it silently fails, and the next time an office IP trips a scenario you get an outage and no idea why.
So I actually tested it. Fed a synthetic failed-login line through the pipeline with one of the whitelisted IPs.
$ cscli explain \
--log "Oct 6 10:32:12 mail sshd[1234]: Failed password for root from 203.0.113.100 port 12345 ssh2" \
--type syslog
line: ...
├ s01-parse
| └ 🟢 crowdsecurity/sshd-logs (+9 ~1)
├ s02-enrich
| ├ 🟢 my/admin-whitelist (~2 [whitelisted])
| └ 🟢 crowdsecurity/whitelists (unchanged)
└-------- parser success, ignored by whitelist (admin + own infrastructure) 🟢
No scenarios ran. The event was dropped at the enrichment stage before anything could see it. That's the whitelist working the way it's supposed to.
The first time I ran this test, the whitelist showed "(unchanged)" instead of "[whitelisted]". Which meant the parser ran, checked, found no match, and let the event through. The IP wasn't covered. I'd missed one of the entries. Catching that before it became a real problem was the best fifteen minutes I spent on this whole thing.
Cleaning up firewalld, which was mostly fine actually
The handoff doc for this server claimed there were something like 3,600 Fail2ban rules sitting in the iredmail firewalld zone. I went in expecting a mess. Actual count was 20.
$ firewall-cmd --zone=iredmail --list-rich-rules | wc -l
20
$ firewall-cmd --zone=iredmail --list-rich-rules | grep -c 'icmp-port-unreachable'
0
Zero Fail2ban rules in firewalld. The doc was wrong. The real Fail2ban rules were all in the iptables chains we'd already deleted. Someone probably wrote the doc from memory or from an assumption about what Fail2ban does.
The 20 rules that were there broke into three groups. Three rate-limit rules on ports 25, 465, 587 (someone added those manually years ago as an anti-flood). One geoblock rule using an ipset. And 16 static drops that had nothing to do with Fail2ban — they were just hardcoded IPs and CIDRs.
Ran a whois on the static drops to figure out what they were. Four of them were known malicious hosting netblocks, the kind of bulletproof VPS providers that spam and scan constantly. Those I kept. Two of them turned out to be a /16 belonging to LinkedIn and a /16 belonging to a Taiwanese mobile carrier. Someone at some point blocked an entire mobile carrier's subscriber range because one spammer used one IP from it. That silently drops mail from legitimate users with no bounce and no log entry. Also a /21 from a Malaysian mobile carrier, same story.
Removed those three. Left the four malicious hosting netblocks alone, they're doing their job. Removed nine individual /32 IPs that hadn't been touched in years — nothing against them, they're just cold, and CrowdSec will catch them if they come back.
Down to 8 rules. Three rate-limits, one geoblock, four known-bad netblocks. Every rule is deliberate now.
What it looks like 24 hours later
CrowdSec on its own for a full day. The alerts:
postscreen-rbl caught 2,306 spam sources. That one's interesting because Fail2ban never had DNSBL logic at all, so all of that was just... going through before, presumably getting rejected by Postfix but not being logged as a Fail2ban concern. Now it's tracked.
postfix-non-smtp-command, 335. postfix-spam, 203. ssh-bf, 152. ssh-slow-bf, 149. dovecot-spam, 73. ssh-bf_user-enum, 59. ssh-slow-bf_user-enum, 48. HTTP probes and CVE attempts, another couple dozen. About 3,500 alerts total in 24 hours.
The Romanian netblock that got 4-hour bans the first day came back and got 144-hour bans the second. Same nine IPs, same coordinated attack pattern, now banned for six days.
Also caught the Chinese mobile IP that had been brute-forcing SSH from outside. It showed up in the login banner ("25 failed login attempts since last successful login") and CrowdSec had already banned it. That's a nice feeling.
What I actually learned
The thing that finally clicked for me is that fail2ban-client banned is not a reliable indicator of what's blocked. The internal state and the kernel state can drift. On a system with mixed iptables/nftables backends, orphaned rules happen. If you're debugging a "this IP is blocked but shouldn't be" problem, check the kernel, not the client.
Second thing: when a problem is "works everywhere except one place," test each port independently. The fact that IMAP worked and SMTP didn't was the entire clue. If I'd only tested 993 I'd have spent hours on the wrong layer.
Third: verify your whitelist actually whitelists. Loading the file and reloading the service is not the same as confirming the entries match. Run a test event through the pipeline. cscli explain is right there, use it.
Fourth: documentation lies. That "3,600 rules" number cost me time I didn't need to spend. Measure the live system.
Fifth, and this is the one that matters: the real fix wasn't tuning Fail2ban. I could have added the office IP to ignoreip, written a cron job to periodically flush orphaned kernel rules, tightened thresholds. All of those are band-aids that leave the dual-firewall architecture in place. The actual fix was removing the second ban system entirely so this class of bug can't happen. Fail2ban was the source of the orphaned rules. Fail2ban is gone. Bug class eliminated.
The single nftables table CrowdSec maintains is atomic. Adding a decision updates the kernel instantly. Removing a decision updates it instantly. There's no separate state to fall out of sync. There's one whitelist instead of two. When something goes wrong on this server in six months, the investigation will be minutes, not days.
That's the whole point of the exercise. Not "fix the thing" but "remove the conditions that let the thing happen." Took two days. Worth it.
If you've hit something similar — IPs you thought were unblocked that turn out to be blocked, security systems that report one thing while the kernel does another, dual-firewall confusion where you don't know which layer is dropping packets — the Fail2ban to CrowdSec move is worth the effort. The migration itself is a few hours if you plan it. The time savings on the next debugging session are much bigger.
IPs in this post use the RFC 5737 documentation ranges and RFC 1918 internal ranges. Domains and organizational details have been changed. All technical sequences match what was actually done.