Claude Opus 5.5, a Popular Humanizer, and SEOZilla
100 articles SEOZilla shipped to real customers, re-written from the same brief by Claude Opus 5.5 — then run through the open-source Humanizer skill as well. All three versions scored by ZeroGPT, the independent AI detector.
Benchmark run September 2026 · ZeroGPT "fakePercentage" · lower = more human
Raw Claude Opus 5.5
45.7%
average AI-detected
Opus 5.5 + Humanizer
42.2%
average AI-detected
SEOZilla Humanised
21.1%
average AI-detected
Free site analysis · 3-day free trial for articles
How We Ran It (So You Can Check Our Work)
No cherry-picking. We took the 100 most recent English-language articles our engine generated for real customer projects, in order, and gave Claude the exact brief a normal person would type.
Same brief, all three
The exact title, target keywords, and word count SEOZilla was given — the newest 100 English articles in a row, no skipping.
A fresh Claude Opus 5.5
Each draft came from a brand-new Claude Opus 5.5 session with no extra instructions — no humanisation hints, no style tricks, first answer kept.
Then the Humanizer skill
A second fresh Opus 5.5 session ran each draft through the open-source Humanizer skill (v3.1.0), exactly as its README describes.
One judge: ZeroGPT
Every version was scored through the identical ZeroGPT pipeline we use in production. Lower score = reads more human.
The exact prompt Claude got (example)
"Claude, I need a SEO article of about 2400 words, titled ‘…’, targeting these keywords: … Write it in markdown."
The Results
First the big picture, then every individual run.
Where the 100 Articles Landed
Number of articles per ZeroGPT score band. SEOZilla clusters in the human zone; raw and humanized Opus both spread into AI-detected territory.
0–20%
reads human
20–40%
mostly human
40–60%
mixed signals
60–80%
likely AI
80–100%
AI generated
Mean score
45.7%·42.2%·21.1%
Median score
45.5%·41.2%·19.2%
Every Single Run
All 100 runs, sorted by SEOZilla's margin over the humanized Opus draft.
Bar length = ZeroGPT AI-detection score (0–100%, lower is better). Differences under 5 points are within detector noise and counted as ties.
The Scoreboard
SEOZilla vs raw Claude Opus 5.5
84
SEOZilla wins
14
ties (within 5 pts)
2
Opus wins
SEOZilla vs Opus 5.5 + Humanizer
85
SEOZilla wins
10
ties (within 5 pts)
5
Humanizer wins
The Stat That Should Worry You
An article that scores 50%+ reads as mostly AI — exactly the kind of content Google and ChatGPT have been deranking. Each square below is one of the 100 articles; red means it crossed that line.
Raw Claude Opus 5.5 drafts
40 of 100
scored 50%+ AI-detected
Opus 5.5 + Humanizer
32 of 100
scored 50%+ AI-detected
SEOZilla articles
1 of 100
scored 50%+ AI-detected
"Just Run It Through a Humanizer" Doesn't Cut It
The Humanizer skill is a popular open-source prompt that rewrites the stylistic tells Wikipedia editors use to spot AI writing — "not X but Y" contrasts, one-line closers, bold labels, em dashes. It makes the prose cleaner. It barely moves an AI detector.
−3.5 pts
average improvement over raw Opus
36 / 5
drafts it helped / made worse (5+ pts)
5
articles where it beat SEOZilla
To its credit, the Humanizer's own README says it plainly: getting past AI detectors is not its goal, and detectors still flag most of its output. It edits for human readers — which is exactly why it isn't a substitute for a humanisation pipeline that is verified against detectors.
Claude Is a Brilliant Writer. That's Not the Problem.
We use frontier AI models inside SEOZilla too. The difference is what happens after the first draft: every article runs through our proprietary humanisation pipeline and is scored against AI detectors before it ships.
A raw model, prompted once
- One draft, straight to you — no detector ever sees it
- 40 of 100 drafts scored 50%+ AI-detected
- Only 1 of 100 scored under 10%
A raw model + a style-cleanup prompt
- Removes surface tells like dashes, bold labels, and stock phrases
- Barely moves what detectors actually measure
- 32 of 100 still scored 50%+ AI-detected
SEOZilla's humanisation pipeline
- Multiple humanised variants, detector-scored, best one ships
- 21 of 100 articles scored under 10% AI-detected
- Plus real keyword research, internal links, images, and auto-publishing
The fine print (because benchmarks without it are marketing fiction)
- All 300 scores were run fresh, in the same session, through the identical pipeline — SEOZilla articles exactly as published (summary box included), Opus drafts raw and humanized. Failed detector calls were retried, never counted as a score.
- SEOZilla serves the better of several detector-scored humanised variants, while Opus got one raw shot — that head start is the product, not a flaw in the test.
- The Humanizer skill is github.com/blader/humanizer, version 3.1.0, applied in its own file mode by a fresh Claude Opus 5.5 session. It is built to improve readability, not to evade detectors, and says so.
- Opus drafts ran about 24% longer than their target word counts; longer Opus drafts actually scored slightly lower, so length doesn't flatter SEOZilla here.
- Single-article ZeroGPT scores wobble by a few points; we treat anything within 5 points as a tie. ZeroGPT is one detector, and the sample leans toward a handful of active customer niches.
- Claude model: claude-opus-5-5, September 2026. Our earlier run against Claude Fable 5 is here.
See What the Pipeline Does to Your Content
Enter your website to get a free site analysis, then try article generation free for 3 days — every article comes with its AI-detection score, so you can verify it yourself.
Free site analysis · 3-day free trial for articles