A real benchmark — 100 articles, 300 AI-detection scores, every result published

Claude Opus 5.5, a Popular Humanizer, and SEOZilla

100 articles SEOZilla shipped to real customers, re-written from the same brief by Claude Opus 5.5 — then run through the open-source Humanizer skill as well. All three versions scored by ZeroGPT, the independent AI detector.

Benchmark run September 2026 · ZeroGPT "fakePercentage" · lower = more human

Raw Claude Opus 5.5

45.7%

average AI-detected

Opus 5.5 + Humanizer

42.2%

average AI-detected

SEOZilla Humanised

21.1%

average AI-detected

Free site analysis · 3-day free trial for articles

How We Ran It (So You Can Check Our Work)

No cherry-picking. We took the 100 most recent English-language articles our engine generated for real customer projects, in order, and gave Claude the exact brief a normal person would type.

1

Same brief, all three

The exact title, target keywords, and word count SEOZilla was given — the newest 100 English articles in a row, no skipping.

2

A fresh Claude Opus 5.5

Each draft came from a brand-new Claude Opus 5.5 session with no extra instructions — no humanisation hints, no style tricks, first answer kept.

3

Then the Humanizer skill

A second fresh Opus 5.5 session ran each draft through the open-source Humanizer skill (v3.1.0), exactly as its README describes.

4

One judge: ZeroGPT

Every version was scored through the identical ZeroGPT pipeline we use in production. Lower score = reads more human.

The exact prompt Claude got (example)

"Claude, I need a SEO article of about 2400 words, titled ‘…’, targeting these keywords: … Write it in markdown."

The Results

First the big picture, then every individual run.

Raw Claude Opus 5.5Opus 5.5 + HumanizerSEOZilla

Where the 100 Articles Landed

Number of articles per ZeroGPT score band. SEOZilla clusters in the human zone; raw and humanized Opus both spread into AI-detected territory.

7
8
55
31
40
38
40
34
7
18
14
0
4
4
0

0–20%

reads human

20–40%

mostly human

40–60%

mixed signals

60–80%

likely AI

80–100%

AI generated

Mean score

45.7%·42.2%·21.1%

Median score

45.5%·41.2%·19.2%

Every Single Run

All 100 runs, sorted by SEOZilla's margin over the humanized Opus draft.

Watch anatomy explainerSEOZilla by 54.0
Habit science mythsSEOZilla by 53.6
Stress supplement explainerSEOZilla by 51.8
Content personalisation guideSEOZilla by 48.9
Fermented food beginner guideSEOZilla by 46.9
Watch crystal comparisonSEOZilla by 46.8
Travel watch features guideSEOZilla by 43.8
Wellness gummies roundupSEOZilla by 43.4
Tactical combat explainerSEOZilla by 42.2
WordPress navigation menu guideSEOZilla by 42.0
Habit stacking guideSEOZilla by 41.6
Game clothing template guideSEOZilla by 40.8
Appetite supplement explainerSEOZilla by 40.3
Online marketing course reviewSEOZilla by 40.2
Personalisation engine explainerSEOZilla by 40.0
SEO strategy tools roundupSEOZilla by 39.2
PC game release roundupSEOZilla by 38.6
SEO writing tools roundupSEOZilla by 38.0
Co-op team setup guideSEOZilla by 36.7
SEO KPIs explainerSEOZilla by 36.2

Bar length = ZeroGPT AI-detection score (0–100%, lower is better). Differences under 5 points are within detector noise and counted as ties.

The Scoreboard

SEOZilla vs raw Claude Opus 5.5

84

SEOZilla wins

14

ties (within 5 pts)

2

Opus wins

SEOZilla vs Opus 5.5 + Humanizer

85

SEOZilla wins

10

ties (within 5 pts)

5

Humanizer wins

The Stat That Should Worry You

An article that scores 50%+ reads as mostly AI — exactly the kind of content Google and ChatGPT have been deranking. Each square below is one of the 100 articles; red means it crossed that line.

Raw Claude Opus 5.5 drafts

40 of 100

scored 50%+ AI-detected

Opus 5.5 + Humanizer

32 of 100

scored 50%+ AI-detected

SEOZilla articles

1 of 100

scored 50%+ AI-detected

"Just Run It Through a Humanizer" Doesn't Cut It

The Humanizer skill is a popular open-source prompt that rewrites the stylistic tells Wikipedia editors use to spot AI writing — "not X but Y" contrasts, one-line closers, bold labels, em dashes. It makes the prose cleaner. It barely moves an AI detector.

−3.5 pts

average improvement over raw Opus

36 / 5

drafts it helped / made worse (5+ pts)

5

articles where it beat SEOZilla

To its credit, the Humanizer's own README says it plainly: getting past AI detectors is not its goal, and detectors still flag most of its output. It edits for human readers — which is exactly why it isn't a substitute for a humanisation pipeline that is verified against detectors.

Claude Is a Brilliant Writer. That's Not the Problem.

We use frontier AI models inside SEOZilla too. The difference is what happens after the first draft: every article runs through our proprietary humanisation pipeline and is scored against AI detectors before it ships.

A raw model, prompted once

  • One draft, straight to you — no detector ever sees it
  • 40 of 100 drafts scored 50%+ AI-detected
  • Only 1 of 100 scored under 10%

A raw model + a style-cleanup prompt

  • Removes surface tells like dashes, bold labels, and stock phrases
  • Barely moves what detectors actually measure
  • 32 of 100 still scored 50%+ AI-detected

SEOZilla's humanisation pipeline

  • Multiple humanised variants, detector-scored, best one ships
  • 21 of 100 articles scored under 10% AI-detected
  • Plus real keyword research, internal links, images, and auto-publishing

The fine print (because benchmarks without it are marketing fiction)

  • All 300 scores were run fresh, in the same session, through the identical pipeline — SEOZilla articles exactly as published (summary box included), Opus drafts raw and humanized. Failed detector calls were retried, never counted as a score.
  • SEOZilla serves the better of several detector-scored humanised variants, while Opus got one raw shot — that head start is the product, not a flaw in the test.
  • The Humanizer skill is github.com/blader/humanizer, version 3.1.0, applied in its own file mode by a fresh Claude Opus 5.5 session. It is built to improve readability, not to evade detectors, and says so.
  • Opus drafts ran about 24% longer than their target word counts; longer Opus drafts actually scored slightly lower, so length doesn't flatter SEOZilla here.
  • Single-article ZeroGPT scores wobble by a few points; we treat anything within 5 points as a tie. ZeroGPT is one detector, and the sample leans toward a handful of active customer niches.
  • Claude model: claude-opus-5-5, September 2026. Our earlier run against Claude Fable 5 is here.

See What the Pipeline Does to Your Content

Enter your website to get a free site analysis, then try article generation free for 3 days — every article comes with its AI-detection score, so you can verify it yourself.

Free site analysis · 3-day free trial for articles