Skip to content

When a Botnet Took Down Four WordPress Sites: What We Found and Changed in Trellis

• By Jasper Frumau DevOps

On 3 October 2026 our own site went down for 30 minutes. updown.io reported a 30-second response timeout from every probe location, from 05:49 to 06:18 CEST. The TLS handshake worked, the certificate was valid, and plain HTTP redirected as usual. Requests simply never came back.

The server is one Hetzner machine running Trellis, and it hosts four sites: this one, a theme demo network, a news site and a separate app. This post is the post-mortem: what we assumed, what the logs actually showed, and the changes we made in Trellis so one noisy site can no longer take the others down.

Quick Summary: A flood from about 32,800 distinct IPs in one hour hit the demo site’s login page and an uncacheable faceted-navigation URL pattern. All four sites shared one PHP-FPM pool of 40 workers, so the pool filled and every site timed out. IP blocklists cannot stop a flood spread over tens of thousands of addresses. What helps is isolating pools per site, closing the URL pattern, and throttling wp-login.php by site rather than by IP. All three are Trellis changes you can provision. After applying them, load on the server went from about 41 to about 0.5.

What’s Covered

Our First Guess Was Wrong

We had just renewed a wildcard Let’s Encrypt certificate for one of the other sites, so the first suspect was the certificate work. That was quickly ruled out: the certificate on the affected site had been issued two days earlier and was valid until December.

The server itself told the real story. Load average was around 41 on a four-core machine, nginx and PHP-FPM were both running, and the PHP-FPM log held repeated warnings that the pool had reached pm.max_children (40). Nginx accepted connections fine, but every request queued behind PHP workers that were all busy.

Our second guess was also wrong. A quick tally of the busiest IPs pointed at one address that had sent about 4,700 requests to the news site’s homepage, so we blocked it along with a handful of brute-forcers. That felt productive, and load did fall shortly after. But it was the wrong cause, and we cannot honestly claim the block was what brought the site back.

What the Logs Actually Showed

That single IP was making roughly 30 requests a minute, mostly 404 scans. Annoying, but nowhere near enough to saturate the server. When we counted requests per site per hour instead of per IP, the picture changed completely.

Hour (CEST)Demo site requestsEverything else
03:0012,657under 200 per site
04:0013,202under 600 per site
05:0035,366under 400 per site

In the 05:00 hour the demo site received about 10 requests a second from 32,817 distinct IPs. No single address sent more than roughly 130 requests. Two URL patterns accounted for almost all of it:

  • About 17,100 requests to wp-login.php?redirect_to=…
  • About 16,600 requests to blog-archive URLs with combinations of ?query-0-page=…&query-1-page=…, a faceted-navigation pattern that generates a unique URL for every combination

Neither can be served from the page cache. The login page is excluded by design, and each unique query-string combination is a cache miss that boots WordPress. In that one hour 4,259 requests got a 502 and about 13,000 were abandoned by the client before the server answered.

Why an IP Blocklist Was Not the Fix

We keep a reviewed deny list in Trellis, with every entry checked against AbuseIPDB. It is the right tool for scanners and brute-forcers that come from a handful of addresses. Against 32,000 sources it does nothing: the next hour’s attackers are different addresses, and none of the ones we blocked were among the ones hitting the demo site.

The deeper problem was architectural. Trellis, by default, runs one PHP-FPM pool for every site on the server. A flood aimed at the least important site, a demo network, consumed the workers that serve the production sites too. The blast radius of any single site was the whole server.

The Three Trellis Changes

1. A dedicated PHP-FPM pool for the demo site

We added an optional php_fpm block to a site’s entry in wordpress_sites.yml. Sites that define it get their own pool and socket; sites that do not keep using the shared one. The demo site now has eight workers, and the shared pool dropped from 40 to 32, so total capacity (and therefore worst-case memory) is unchanged. A flood on the demo site can now only exhaust its own eight workers.

For the developers: the Trellis wiring

The site entry opts in; the nginx site template points fastcgi_pass at that site’s socket, falling back to the shared one.

# group_vars/production/wordpress_sites.yml
demo.imagewize.com:
  php_fpm:
    pool: wordpress-demo
    max_children: 8
    start_servers: 2
    min_spare_servers: 1
    max_spare_servers: 4

# roles/wordpress-setup/templates/wordpress-site.conf.j2
fastcgi_pass unix:/var/run/php-fpm-{{ item.value.php_fpm.pool | default('wordpress') }}.sock;

# roles/wordpress-setup/tasks/main.yml: render one pool file per opted-in site
- name: Create dedicated php-fpm pools for sites that define php_fpm.pool
  template:
    src: php-fpm-pool-site.conf.j2
    dest: /etc/php/{{ php_version }}/fpm/pool.d/{{ item.value.php_fpm.pool }}.conf
  loop: "{{ wordpress_sites | dict2items | selectattr('value.php_fpm.pool', 'defined') | list }}"
  notify: reload php-fpm

2. Close the crawler trap

Real visitors paginate one query loop at a time, so a URL carrying two or more query-N-page parameters is never legitimate. We return 410 Gone for those, and for the ?cst variant, from a per-site nginx include. A single ?query-0-page=2 still works. We had spotted the same pattern in August and blocked three IPs then; this time the fix is on the URL pattern, not the sender.

For the developers: the nginx include
# nginx-includes/demo.imagewize.com/block-crawler-trap.conf.j2
if ($args ~* "query-[0-9]+-page=[^&]*&.*query-[0-9]+-page=") {
    return 410;
}
if ($args ~* "(^|&)cst($|=|&)") {
    return 410;
}

3. Throttle wp-login.php by site, not just by IP

A per-IP rate limit is useless when each IP only sends a handful of requests. So we added a site-wide ceiling on the login page (2 requests a second, with a small burst) alongside a per-IP limit (20 a minute). The trick is two nginx map blocks that produce a limit key only for login URLs; every other URL maps to an empty key, and nginx does not count empty keys, so nothing else is affected. Over the limit, nginx answers 429.

For the developers: zones and the server-level include

Zones and maps live in the http context (nginx.conf.j2); the limits are applied through an include shared by all sites.

# http context
map $uri $wplogin_site_key {
    ~^/(wp/)?wp-login\.php$ $server_name;
    default "";
}
map $uri $wplogin_ip_key {
    ~^/(wp/)?wp-login\.php$ $binary_remote_addr;
    default "";
}
limit_req_zone $wplogin_site_key zone=wplogin_site:1m rate=2r/s;
limit_req_zone $wplogin_ip_key   zone=wplogin_ip:10m  rate=20r/m;

# nginx-includes/all/rate-limit-login.conf.j2 (included in every server block)
limit_req zone=wplogin_site burst=10 nodelay;
limit_req zone=wplogin_ip   burst=5  nodelay;
limit_req_status 429;

Provision with trellis provision --tags nginx,wordpress-setup production. We previewed every change first with ansible-playbook server.yml -e env=production --tags nginx,wordpress-setup --check --diff.

One thing we considered and deliberately skipped: caching 404 responses. WordPress sends no-cache headers on 404s, so nginx will not cache them unless you tell it to ignore upstream headers, which is risky to do globally.

Verifying It Works

After provisioning we tested each change against production rather than trusting the config:

CheckResult
All four sitesHTTP 200
Two query-N-page parameters, or ?cst410
A single ?query-0-page=2200 (real pagination untouched)
25 rapid requests to wp-login.php15 served, 10 answered 429
PHP-FPM poolsShared pool plus a separate wordpress-demo pool, each with its own socket
Server loadAbout 0.5, down from about 41

What we have not verified: how the new limits behave under a real 30,000-IP flood. We tested the mechanisms individually, not the attack. We will report back if there is a next time.

Detection Was Slower Than the Outage

The updown.io timeline showed the first failure at 03:49 UTC and the confirmed-down alert at 03:59 UTC. We were told ten minutes late, because the check interval was about ten minutes and the alert waited for several probe locations to agree. The outage lasted 30 minutes, so a third of it passed before anyone knew.

The fix there is plain configuration: a one-minute check interval, alerting after two failed checks, and separate checks for every site on the server, since a shared-server failure takes them all down together.

Why We Did Not Reach for a CDN

The usual advice is to put a large US CDN or WAF in front of everything. We chose to see how far the server itself could go first, partly because we prefer to keep traffic and data inside EU providers. A flood of 10 requests a second is well within what a properly isolated, properly throttled server can absorb. If it ever is not, the next step is an EU option such as Bunny.net, CDN77, Gcore or Myra rather than a default to the biggest name.

What We Would Tell Another Trellis Owner

  • Count requests per site before per IP. Our busiest single IP was not the problem. The busiest site was.
  • Do not let one site share workers with the rest. Give low-value sites a small pool of their own.
  • Rate-limit by site as well as by IP for the pages that must never be cached, like the login page.
  • Treat a deny list as a tool for repeat offenders, not a defence against distributed floods.
  • Check your monitoring interval. If the check is slower than your tolerance for downtime, it is not monitoring.

Done Managing Your Own Server?

We offer managed WordPress hosting built on Trellis, with per-site PHP-FPM pools, Nginx rate limiting, FastCGI caching and Redis object cache, automated deployments via Ansible, and Bedrock structure on Hetzner EU. No shared hosting, no page builders, no surprises.

  • Trellis + Bedrock on Hetzner EU (Frankfurt / Helsinki)
  • Nginx + FastCGI caching + Redis object cache
  • Automated deployments via Ansible, SSL via Let’s Encrypt
  • From €49/month, or €65/hour for one-off server work

Leave a Reply

Your email address will not be published.