WordPresssecurityfleet management

Three threats, all marked High. Two were a backdoor, one was a plugin update.

At 01:53 on a Wednesday morning, on a WordPress site I look after, six plugins were deactivated in thirty seconds. They were the security plugins: the scanner, the firewall, the backup, the spam filter. Nothing else on the site was touched.

That is the textbook opening move, and on this occasion it achieved nothing at all.

At 09:58 the same morning, the malware scanner found the attacker's dropper. At 10:00 it emailed us. The email had the full path in it:

/srv/htdocs/wp-content/plugins/contact_1788357453/page_template_1788357454.php

That is the exact file I quarantined. I quarantined it seven days later, after a human visited the site, got shown a fake captcha telling them to paste a command into the Windows Run dialog, and mentioned it.

The scanner was not blinded. The alert was not lost. The file path was in an inbox for a week.

And that is the flattering version of this story, because it was not the first email.

The email that actually mattered arrived seven weeks earlier

The backdoor did not land in September. The first malicious file on that site was written on the thirteenth of June, an mu-plugin that injected the loader into every page and quietly built a list of every administrator IP address so it could hide from us specifically.

On the fifteenth of July, thirty-two days after that file appeared, the scanner emailed. Here is that email, in full, with only the site name removed:

Your site may be at risk
Jetpack Scan found 3 threats on [site]

High
The database table wp_options contains malicious code
Threat found: database_malware_etherhide_005

High
The database table wp_options contains malicious code
Threat found: database_malware_etherhide_005

High
Vulnerable Plugin: gravityforms (version 2.10.3)
Vulnerability found in plugin

Read that again, because it is the whole argument in one screenshot.

Two of those three are a live compromise. database_malware_etherhide_005 is the signature for malware that keeps its payload on a blockchain, and the vendor's own description of it says its presence means malware is or was on the site and that you should go and review your files. It was correct. It had found the fingerprint of the June injector.

The third is a plugin that needs updating.

All three are High. All three are in the same email, in the same format, one after another. There is nothing in that message that separates "somebody is inside your server" from "there is a version bump available for a form plugin you run on 240 sites."

I had assumed, before I went looking, that the noise problem was one of volume: hundreds of routine emails and one real one, and the real one gets lost in the pile. That is true, and it is not the worst of it. The routine finding is inside the same email as the compromise, wearing the same severity label. You cannot fix that by reading your mail more carefully.

Nobody actioned it. The real distance from the first correct, delivered alert to the day we found the backdoor is not seven days. It is fifty-six.

And July was the cheap moment. At that point the site had one malicious file on it and no webshells. Everything the attacker added later, the second injector, the self-healing copies hidden in the uploads folder, the two remote-execution backdoors, all of it came after that email had already been sent and read past.

Turning off the plugin did nothing, and that surprised me

I assumed, when I started writing this, that deactivating the scanner is what stopped the alert. That is wrong, and it is worth being specific about why, because the mistake is a common one.

Jetpack Scan does not run on your site. There is no scanning engine in the plugin. The class that produces those findings is documented as handling "fetching of threats from the Scan API", it defines SCAN_API_BASE = '/sites/%d/scan', and it calls out to WordPress.com. The scanning happens on Automattic's infrastructure, and the plugin's job is to fetch the results and draw them in wp-admin.

So deactivating it closed a window. It did not stop anything from looking. The scan ran, found the dropper eight hours after the plugin was switched off, and the email went out ninety seconds after detection.

Which means the interesting failure is not the one I expected to write about.

Two hundred and forty to one

Here is the fleet this arrived into. 292 WordPress sites. Every one of them running Jetpack. 6,321 plugin installations between them, across 841 distinct plugins. Gravity Forms alone is on 240 of those sites.

Now consider what happens when a single vulnerability is published against Gravity Forms. That is not one email. That is up to 240 emails, one per site, each of them titled to say that a site may be at risk and a threat was found.

The webshell generated one email.

Same subject line. Same severity word. Same sender, same shape, same colour of banner. An actual remote code execution backdoor on a live client site is, in the inbox, visually indistinguishable from a routine notice that a plugin on 240 sites has a version bump available.

And as the July email shows, they do not even need to be separate messages. Gravity Forms turned up in the same email as the malware, at the same severity, three lines below it.

This is the part that people who do not run fleets tend to underrate. The problem is not that anyone was careless with the important email. The problem is that the important email had no distinguishing features. Nothing about it, at a glance, separated it from the ones that arrive every week and are correctly ignored every week.

Put a person in front of a stream where the base rate is a couple of hundred routine to one urgent, and give them no way to tell the two apart without opening each one and reading a file path, and they will miss the urgent one. Not sometimes. Reliably. That is not a discipline problem, it is an information problem, and it is the vendor's design that creates it.

Two events that should never share an envelope

"A plugin you run has a published vulnerability, and an update exists" is a maintenance ticket. It is real, it matters, it belongs in a queue, and it can wait until Thursday.

"A file on your server contains a malicious code pattern" is not that. Somebody is already inside. There is nothing to schedule.

Those are different events with different urgency and different responses, and every mainstream WordPress security product I have used delivers them through the same channel in the same format. Once you have more than a handful of sites, that single decision does most of the damage. The second category is drowned by the first, and the drowning is structural rather than accidental.

Worse, the routine category scales with your fleet and the urgent one does not. Add a hundred sites and your maintenance noise goes up by a hundred sites' worth. The number of actual compromises does not move. So the ratio gets worse precisely as the fleet gets big enough for the stakes to be real.

The honest bit about two-factor

The obvious objection is that none of this matters if the attacker cannot log in, and two-factor authentication is how you stop that. Correct. It is the real fix for the first step, and I am not going to pretend otherwise.

I am also not going to pretend rolling it out across a few hundred WordPress sites is solved, because the obstacles are boring rather than technical:

  • Every site has its own user table. There is no central identity. Three hundred sites is three hundred separate lists of people, each with its own idea of who is an administrator.
  • Half the accounts are not yours. Clients, their staff, a marketing contractor, the agency that built the site in 2019 and still has a login. You can require a second factor for your own team tomorrow. Requiring it of a client's office manager is a negotiation, not a config change.
  • The enforcement is itself a plugin, which anyone who gets in can deactivate.
  • Recovery lands on you, at whatever hour somebody changes phones.

So: do it, prioritise it, start with the accounts that appear on the most sites. And assume it will be incomplete for a long time, which means the day someone gets in is not hypothetical and the alerting path has to work.

What I think the fix actually is

Not better detection. The detection was perfect. It found a dropper the same morning it landed and told us where it lived.

The fix is that findings need to arrive already sorted, and the sorting has to happen somewhere that knows the whole fleet:

  • Separate the classes hard. Malicious file present and vulnerable version installed are different products, not two severities of one product. They should not arrive by the same route.
  • Deduplicate by advisory, not by site. One Gravity Forms CVE is one thing that happened, not 240 things. It should produce one item that names 240 sites.
  • Anything in the compromise class gets an owner and a clock, and stays visible until somebody closes it with a reason. An email is not a queue. It has no state, no assignee, and no way to tell you it was ignored.

That is the direction Sentry is built in, and I want to be accurate about where it currently stands: the outside-looking-in collection is real and running, and the alert routing described above is designed and not finished. Saying otherwise would be exactly the kind of claim this note exists to complain about.

Where I got this wrong

Two places, and they were both mine.

Three weeks before the compromise I ran a malware sweep across the whole fleet, including this site. It came back clean. The sweep was not broken; it searched for the indicators from the previous incident I had dealt with, and this attacker used none of them. A signature list copied forward from the last incident can only ever catch the last incident.

I also wrote the first published version of this note saying the alert sat unread for seven days. It was fifty-six. I had the September email in front of me and did not think to ask whether there had been an earlier one, which is the same failure the note is about: I looked at the loudest signal and stopped.

And the first draft of this note argued that the attacker blinded the scanner by switching it off. I had two true timestamps, thirty seconds apart, and I built a tidy causal story between them without checking whether the scanner was even running on the site. It was not. The alert had already been sent, and the answer was sitting in an inbox the whole time I was writing about why no alert arrived.

Ten minutes, if you run more than a few WordPress sites

Search your mail for the last six months of security notifications from whatever you run. Then answer one question: if one of them had been real, what about it would have looked different?

If the answer is "the wording inside, once I opened it," then your alerting is a filing system rather than an alarm, and it will fail in exactly the way described here. Mine did.

The second read, on any site you care about: wp option get recently_activated --format=json. WordPress records the moment every plugin was deactivated and malware does not clean it up. A cluster of security plugins sharing a timestamp nobody on your team will claim is a compromise until proven otherwise. That is where the 01:53 came from.

We write these when something in the work turns out to be worth writing down. If you're hitting the thing described above and want a second set of eyes on it, tell us what's not working.

← All notes