Spam Blog (Splog): Why Real Sites Trip the Same Signals

Your traffic fell after a core update, nobody at Google will tell you why, and somewhere in the diagnosis a consultant used the word splog. You didn’t scrape anything. You didn’t buy a dropped domain. Every post carries your name. The definition of a spam blog is the easy part. The harder question is whether you’d recognize one if you were quietly running it.

Two reactions follow, and both cost something. The first treats splog as a slur reserved for obvious garbage, keeps the publishing schedule exactly as it was, and takes a second hit on the next update. The second panics: comments off, half the archive set to noindex, pages deleted that were still earning, and no way afterward to tell which change did what.

There is no splog detector. There are spam policies, and they describe behavior rather than intent.

The word describes a production method, not a topic

Splog is a portmanteau of spam and blog. Wikipedia’s entry on spam blogs traces the term’s public use to mid-August 2005, when Mark Cuban used it to name the wave of auto-created blogs then filling free hosting platforms. The scale was not theoretical. That October, a coordinated run produced more than 13,000 splogs on Blogger’s BlogSpot service across a single weekend.

One detail gets repeated wrong almost everywhere. WordPress.com wasn’t part of that wave. It entered private beta on August 8, 2005 and only opened to the public on November 21, so the platform most people now associate with blogging didn’t exist as an open service while the term was being coined. The original splog problem was a Blogger problem.

The traits have barely changed since:

  • Content is scraped, spun, machine-generated, or pulled in from someone else’s feed
  • The site exists to carry display ads, affiliate links, or links pointed somewhere else
  • There is no identifiable author, business entity, or contact route
  • Internal links run to exact-match anchors rather than to genuinely related pages
  • Publishing volume has no plausible relationship to the number of people producing it

Topic has nothing to do with it. A splog about fly fishing and a splog about tax software are the same object wearing different keywords.

Method, not subject.

Google’s spam policies name behaviors, not labels

Google has never published a splog detector, and the system people reach for when they describe one has been retired. The helpful content system was folded into core ranking in March 2024. Google’s own ranking systems guide now lists it under retired systems:

“In March 2024, it evolved and became part of our core ranking systems.”

The same documentation adds that there’s no longer one signal or system used to do this.

So any advice built around “the Helpful Content signal” as a live, separate thing is describing something Google formally shut down.

What Google does name is SpamBrain, which it describes as one among a range of spam detection systems it employs. Around that sits the published spam policies for Google Web Search, and six entries on that list are what get applied to a blog accused of being a splog. The current names matter, because most guides are still quoting the 2023 ones.

  1. Scaled content abuse. High-volume, low-value pages made mainly to game rankings. The policy applies “no matter whether content is produced through automation, human efforts, or some combination of human and automated processes.”
  2. Expired domain abuse. Buying a dropped domain for its residual authority and refilling it with unrelated content. Google’s own illustration is a casino site standing on a former medical domain.
  3. Site reputation abuse. Third-party commercial content published on a host domain mainly because of that host’s established ranking signals.
  4. Scraping. Republishing other people’s articles with light rewrites. The name most guides still use is “scraped content.”
  5. Thin affiliation. Affiliate pages with no original analysis behind the offer. The name most guides still use is “thin affiliate content.”
  6. Doorway abuse. Google’s listed example is multiple domains or pages targeted at specific regions or cities that funnel users to one page. The name most guides still use is “doorway pages.”

The thin affiliation entry carries a caveat most summaries drop. Google states plainly that not every site participating in an affiliate program is a thin affiliate, and that the ones that aren’t add original product reviews, rigorous testing and ratings. The violation is the empty page, not the link on it.

Site reputation abuse is the entry with a documented timeline. Google announced it on March 5, 2024, made it effective May 5, and began issuing manual actions the following day. It has never named a penalized site. Forbes Advisor, CNN Underscored and WSJ Buy Side were named by SEO trade press rather than by Google, which is worth holding onto before you treat any of them as a documented case.

The mechanics of that policy, and what the affected sections actually recovered, sit in the breakdown of why parasite SEO stopped paying.

One more correction, because it circulates constantly. The August 2024 core update is widely described as an AI crackdown. Google’s announcement says nothing about AI, AI detection, or information gain. Its stated purpose was “showing more content that people find genuinely useful and less content that feels like it was made just to perform well on Search,” alongside a commitment to surface small or independent sites creating useful, original content.

Google’s actual position on AI sits elsewhere and it’s narrower than the rumor: automation, including generative AI, is spam when the primary purpose is manipulating rankings in Search.

The tool isn’t the test.

The signals that catch legitimate sites

Pattern matching doesn’t read intent. A site with a real author, a real business, and a real archive can still produce the exact shapes these policies describe. Six ways it happens.

Comparison sheet pairing six ordinary publishing habits with the Google spam policy entry each one matches: model drafts given a light cleanup under scaled content abuse, an autoblogging plugin pulling RSS under scraping, 10,000 templated city pages under doorway abuse, an unfiltered comment form under user-generated spam, 4,000 single-post tags under thin content, and reformatted vendor pages under thin affiliation.
None of the six needs bad intent, which is why an audit that looks for motive comes back empty.

Unreviewed AI drafts. Publishing model output with a light cleanup pass. The scaled content abuse policy is explicit that human involvement doesn’t exempt the result, so “a person went over it” is not the defense most people assume it is.

Auto-imported feeds. Autoblogging plugins are still live and still installed. WP Content Pilot describes itself as an autoblogging suite with content spinning and automatic affiliate link injection, and carries more than 800 active installs. FeedWordPress, at over 9,000 installs, pulls RSS straight into posts. Both do exactly what they advertise, and what they produce is the scraping policy in plugin form.

Template pages with nothing underneath. 10,000 city pages generated from one template with the city name swapped is Google’s own doorway abuse example, almost word for word. Programmatic pages built on genuine per-page data are a different animal, and even those frequently underdeliver: one build shipped 748 pages and returned 13 clicks in 90 days.

Unmoderated comments. An open comment form with no filter fills the bottom of every post with casino links and crypto pitches, and the crawler reads that as your page content. It also has a manual action label of its own, User-generated spam. The fixes are cheap and mostly settings-level, covered in the walkthrough on stopping WordPress comment spam.

Tag archives that outnumber posts. 4,000 indexable tags, most of them holding a single post, generate thousands of near-empty archive pages that map cleanly onto thin content. Prune them.

Affiliate pages with no original work. If every commercial post reads like a reformatted vendor page, thin affiliation is the entry that applies, and Google already published the exit route: original reviews, testing, ratings.

None of these require bad intent, which is why treating splog as a moral category leads people to the wrong audit.

Splog, content farm, and MFA site describe different failures

These three labels get used interchangeably and they point at different objects. The separating questions are what produced the content and what the site is monetizing.

TypeContent sourceHuman involvementPrimary monetization
SplogScraped, spun, or auto-generatedNear zeroDisplay ads, redirects, link placement
Content farmWritten cheaply at volume by peopleMinimal editingDisplay ads, affiliate
MFA (made for advertising)Human-written, shaped around ad slotsModerateProgrammatic display ads
Legitimate blogWritten by someone with domain knowledgeHighMixed: ads, affiliate, services, products
AI-assisted blogModel-drafted, rewritten by a personModerate to highMixed

The row worth staring at is the last one.

An AI-assisted blog running at volume with one editor skimming drafts sits closer to the content farm row than its owner usually believes, and the distance between those two rows is measured in editing time nobody wants to spend.

What keeps a real blog out of the pattern

None of this is about proving you’re human. It’s about making the page carry something a script has no access to, and keeping the site’s shape consistent with the size of the operation behind it.

Put something first-party in every post

Every substantial post should carry at least one thing that can’t be generated: a screenshot from your own account, a number you measured, a quote from someone you spoke to, a configuration you ran and the output it gave. “Studies show” is not first-party evidence.

Neither is a stock photo of a laptop.

Grow publishing volume at a rate you can defend

A three-week-old domain publishing 20 posts a day is a pattern. A site steady at 2 posts a week for a year that suddenly jumps to 40 is also a pattern, and the second one surprises people. Both look automated.

A workable ceiling:

Double your monthly output at most, then hold that level for 60 days before scaling again.

That’s a working rule, not Google’s. What it buys you is a ramp slow enough that a quality problem surfaces in your own review queue before it surfaces in your rankings.

Build a comment stack that holds

Akismet ships inside WordPress core alongside Hello Dolly, but it isn’t active on a fresh install and does nothing until you add an API key. Its Personal plan is name-your-price, which reads as free until you check the terms: Akismet bars that plan from any site carrying ads or affiliate links, and states that noncompliance results in immediate suspension of services without notice. If your blog monetizes, you belong on a paid tier. The current Akismet plans run Personal, Pro, Business and Enterprise, with Pro priced by monthly spam-check volume starting at 500 checks a month on 1 site.

Budget for it.

Antispam Bee, maintained by pluginkollektiv, is the alternative when that pricing or that data flow is a problem. It’s free for both private and commercial projects, needs no API key and no account, works without captchas, and doesn’t send personal information to third-party services. It carries more than 700,000 active installs and was last updated in May 2026.

For higher-volume forms, put Cloudflare Turnstile in front. Cloudflare calls it a smart CAPTCHA alternative that works without showing visitors a CAPTCHA, which is accurate; “invisible” is one of its three widget modes, alongside Managed and Non-interactive, so the flat description oversells the default. The free Turnstile plan covers unlimited challenges across up to 20 widgets, which is more than a single blog will ever need.

Then handle trackbacks. They are almost entirely spam and no configuration improves them.

Turn them off.

Prune thin pages and tag archives

You need an inventory before you need a strategy. WP-CLI gives you one in a line:

wp post list --post_status=publish --format=csv

Pipe that into a spreadsheet, sort by length, and the thin-content list writes itself. Expand what deserves it, noindex the rest, and apply the same test to tag archives: a tag holding fewer than 5 posts should be merged into a broader one or deleted. More of this kind of bulk triage is in the practical WP-CLI command cheat sheet.

Make the site attributable

An author page with a real name and photo, a contact route that reaches a person, and a business entity where one exists. Splogs skip all of it because attribution is the expensive part of running one at scale. For a real site it’s an afternoon.

Treat model output as a draft

Rewrite the opening, replace generic examples with specific ones, cut the tells, and add the first-party detail the model had no access to. Budget 30 to 60 minutes per 2,000 words for that pass and protect the time. Skip it and you’re publishing at splog quality with your byline attached to it.

If Google has already flagged you

A traffic drop on its own doesn’t tell you which system moved, and the diagnosis order matters more than the speed. Work through it like this.

  1. Check for a manual action first. In Search Console, open Security & Manual Actions, then Manual actions. The labels relevant here are “Major spam problems,” “Thin content with little or no added value,” “User-generated spam,” “Site abused with third-party spam,” and “Site reputation abuse.” If a guide tells you to look for “Pure spam,” it’s out of date. That label is no longer on Google’s list.
  2. Audit the archive, not the homepage. Look for posts under 300 words, pages with no internal links pointing at them, pages with 5 or more outbound affiliate links, duplicate meta titles, and spam sitting in comment bodies. The full sequence is in the Google Search Console SEO audit workflow.
  3. Export the comment queue. A spam folder holding 10,000 entries against 50 approved comments is a moderation setting, not a spam wave. Switch on manual approval under Settings, Discussion.
  4. Noindex before you reach for removals. Set the problem posts to noindex in Rank Math or Yoast and let that do the actual work. Search Console’s Removals tool is temporary: Google states a successful request lasts only about six months, and that using the tool alone won’t work. It buys a window. Nothing more.
  5. File a reconsideration request only if you have a manual action. There is nothing to reconsider on an algorithmic demotion, and filing anyway wastes the one channel you have.
Decision path branching from Search Console's Manual actions report: a listed label means a manual action a human reviewer issued, where a reconsideration request naming the problem, the steps and the counts is read by a person; nothing listed means an algorithmic demotion with no reconsideration channel, no notification when it lifts, and no published recovery timeline.
Both routes need the same cleanup. Only one of them comes with somewhere to report it.

When you do file, the shape Google asks for is specific:

Name the exact quality issue, describe the steps you took to fix it, and document the outcome with counts: URLs removed, comments purged, pages rewritten.

A person reads it.

Google states that manual actions are issued when a human reviewer determines a violation, and a request that names neither the issue nor the fix gives that reviewer nothing to act on.

What cleanup won’t fix

Cleaning up removes the reason for a demotion. It doesn’t hand the traffic back, and the gap between those two things is where most recovery plans quietly fall apart.

Google publishes no recovery timeline for an algorithmic demotion, and there’s no notification when one lifts. Any range you’ve read is somebody’s guess.

You’ll be reading a graph and inferring.

A manual action can be revoked while traffic stays flat, because manual actions and core ranking are separate systems. Getting the notice cleared confirms exactly one thing: the notice is cleared.

The positions you lost may not be available to take back. Something moved into them while you were down, and that page is now the incumbent with months of engagement data behind it.

Comment filters cut volume, not motive. Akismet and Antispam Bee act on comments and nothing else. Neither touches thin posts, tag sprawl, or template pages, which is usually the larger half of the problem and always the slower half to fix.

And nothing reports on information gain. No plugin scores whether your page says anything the current top 10 don’t already say. That judgment stays with you, and it’s the one the whole outcome turns on.

What quietly turns a real site into a pattern

Scaling output because the tooling got cheaper. Generation cost collapsed, editing cost didn’t move, and the ratio between those two is close to what the scaled content abuse policy is measuring.

Installing a fourth anti-spam plugin instead of changing a setting. Each one adds queries on every comment submission and they overlap heavily. One filter plus a challenge on the form outperforms a stack, and costs less on every page load.

Trusting an anti-spam plugin you last evaluated years ago. Plugins get abandoned, and a listing closed on WordPress.org stops receiving security fixes while the code keeps running on your site. Check the listing before you check the settings.

Deleting thin posts instead of improving the ones that had an audience. Deletion is faster and it feels decisive. It also discards links, history, and rankings that a rewrite would have kept.

Adding an author box with nothing behind it. A name and a headshot on a page with no bio, no publishing history, and no reachable contact reads as compliance theater to a reviewer who has seen 10,000 of them.

Republishing your own newsletter or syndicated column into the blog at volume. You own the words, so it feels safe. The crawler sees a mirror and grades it like one.

Splog is a 2005 word for a permanent problem

The 2005 version was easy to spot because the machinery was crude: auto-created accounts, scraped feeds, keyword soup, 13,000 sites in a weekend. The machinery got good.

What Google measures now has almost nothing to do with how the text was produced and nearly everything to do with whether the page adds something the results didn’t already carry.

That reframes the work. There’s no checklist that certifies you as not-a-splog, no plugin that returns a verdict, and no report to screenshot for a client. There’s a short list of behaviors any site can produce, a set of policies that describe them in current language, and one judgment about information gain that no tool will make for you.

The honest trade is that first-party work is slower, more expensive, and looks worse on a dashboard through the first year. It’s also the only part of the operation that somebody with an API key and a template can’t reproduce by Friday.

Open your own archive, read 3 posts at random, and ask what a script couldn’t have written.

Tell Google you want more of this.

Add Gatilab as a preferred source

One tap, and this site shows up more often in your own Top Stories, AI Overviews and AI Mode. Remove it any time.