Research
How Nightpress chooses the models that write your books
Last updated: September 29, 2026
Every Nightpress book is written by a model that won its seat. We run candidate models through the same two-stage bake-off — hard technical gates first, then blind prose judging — and publish the results here, numbers and methodology included, so you can see exactly why the default model is the default.
The outcome: three tiers, straight from the data
Seventeen models went through this process in July 2026 and sorted into exactly two real quality levels. The September register-spanning run added a third: one challenger beat the champion at book length, distinctly, and earned a tier above it. Three tiers, because that's what we measured — and all three on one provider, so caching and batch discounts apply to every book.
Premium
Claude Fable 5.1
The only challenger to clear our swap bar. In the register-spanning run it beat the Standard champion distinctly (win rate 0.72, the whole confidence interval above even), took worldbuilding, voice and physicality outright, and wrote every register with no refusals. Twice the token price; cheap cache reads keep a whole book under twice Standard's cost.
Standard — the default
Claude Opus 4.8
The measured champion of the July campaign. In head-to-head blind judging it beat the flagship that cost twice as much, held off every other challenger over full 8-chapter horizons, and is the only model whose AI-tell density improves as a book deepens. The best quality-per-dollar in the field — which is why it's the default, not just an option.
Budget
Claude Sonnet 5
The honest cheap option: about half of Standard's cost, and it writes every register we have thrown at it, including the dark ones the low-cost GPT models refuse outright. Not yet blind-judged at full book length; that validation is in progress and we'll say plainly where it lands. For drafts, experiments, and high-volume work.
Why did the Budget model change? The previous one, GPT-5.6 luna, and its whole model family answered a dark novel's planning step and its chapters with a flat refusal in September 2026. A tier that can't write horror or a violent thriller isn't a tier. And why a premium tier now, after we said no? Because the data changed: the July flagship lost to Standard; the September one beat it. We sell what we measure.
What a book typically costs in API usage
Estimated from a measured full-pipeline run (42 chapters, 81,500 words) replayed at provider list rates (September 23, 2026) with prompt caching and the Batch API discount applied; the low end is a typical review, the high end a long one. Your actual cost depends on length, review depth, and provider pricing at run time — every book shows its live spend as it runs, against a limit you set.
| Book size | Premium (Fable 5.1) | Standard (Opus 4.8) | Budget (Sonnet 5) |
|---|---|---|---|
| Novella — ~40,000 words | ~$25–32 | ~$15–18 | ~$7–9 |
| Novel — ~75,000 words | ~$47–59 | ~$28–34 | ~$13–17 |
Books currently run up to about 90,000 words (45 chapters), the largest size we have written and checked end to end. Longer books open up once we have run one.
Stage 1: Technical gates
Each candidate writes the first three chapters of a standard test novel in sequence — chapter two sees chapter one, chapter three sees both — against a demanding story bible: first-person narration, a specified voice with named verbal tics, and a central motif that has to carry plot weight. Sequential writing matters because it tests sustained voice, not one-off polish.
The gates, in order: spec compliance (does it obey the bible's point of view, tense, and voice? a gorgeous chapter in the wrong voice is a full rewrite), reliability (every chapter returns cleanly), and mechanics (typographic and structural rules). A model that fails a gate is disqualified no matter how good its prose is.
Leaderboard at a glance
Every model tested, one line each — click any column to sort. “Screen vs champ” is the blind-panel win rate against the reigning champion on the same fixture (0.5 = parity; our bar for changing the write model is a confidence interval clearing 0.5). Prices are USD per million tokens as tested. Full context for every number is in the detailed table below.
| # | Craft | Content | |||||||
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.8Anthropic | baseline | baseline | baseline | $5 | $25 | pass | clean | Champion |
| 2 | Anthropic | 0.723 | 75 | 59 | $10 | $50 | pass | clean | Premium |
| 3 | OpenAI | 0.513 | 54 | 64 | $5 | $30 | pass | clean | Fallback |
| 4 | Claude Opus 5Anthropic | 0.500 | — | — | $5 | $25 | pass | clean | Evaluated |
| 5 | Moonshot | 0.488 | 53 | 38 | $3 | $15 | pass | clean | Premium |
| 6 | GPT-5.5OpenAI | 0.432 | — | — | $5 | $30 | pass | clean | Fallback |
| 7 | Claude Fable 5Anthropic | 0.380 | — | — | $10 | $50 | pass | clean | Premium |
| 8 | OpenAI | 0.314 | 28 | 50 | $2 | $10 | pass | clean | Evaluated |
| 9 | Muse GlimmerMeta | 0.181 | — | — | $0.35 | $1.5 | pass | clean | Budget |
| 10 | Tencent HY3Tencent | 0.125 | — | — | free | free | pass | clean | Budget |
| 11 | GPT-5.6 lunaOpenAI | 0.125 | — | — | $1 | $6 | pass | clean | Budget |
| 12 | GLM-5.2Z.ai | 0.111 | — | — | $1.4 | $4.4 | pass | clean | Budget |
| 13 | GPT-5.4 miniOpenAI | 0.111 | — | — | $0.75 | $4.5 | pass | clean | Evaluated |
| 14 | DeepSeek V4 ProDeepSeek | 0.056 | — | — | $0.435 | $0.87 | FAIL | clean | Disqualified |
| 15 | Gemini 3.1 ProGoogle | 0.042 | — | — | $2 | $12 | pass | clean | Budget |
| 16 | Qwen 3.7 MaxAlibaba | 0.042 | — | — | $1.48 | $4.42 | pass | clean | Budget |
| 17 | MAI-Code-1-FlashMicrosoft | 0.042 | — | — | — | — | pass | clean | Evaluated |
| 18 | GPT-5.4 nanoOpenAI | 0.028 | — | — | $0.2 | $1.25 | pass | clean | Evaluated |
| 19 | DeepSeek V4 Flash 0731DeepSeek | 0.014 | — | — | $0.09 | $0.18 | pass | clean | Evaluated |
| 20 | Gemini 3.5 FlashGoogle | pending | — | — | $1.5 | $9 | pass | clean | Budget |
| 21 | GLM-4.7 FlashXZ.ai | NA | — | — | $0.07 | $0.4 | FAIL | flaky | Disqualified |
| 22 | Grok 4.20xAI | NA | — | — | $1.25 | $2.5 | FAIL | clean | Disqualified |
| 23 | Llama 3.3 70BMeta | NA | — | — | $0.04 | $0.12 | FAIL | flaky | Disqualified |
| 24 | MiniMax M3MiniMax | NA | — | — | — | — | pass | flaky | Disqualified |
| 25 | GPT-5.6 terraOpenAI | pending | — | — | $1 | $6 | pass | clean | Evaluated |
| 26 | Claude Sonnet 5Anthropic | NA | — | — | $2 | $10 | pass | flaky | Disqualified |
| 27 | Claude Opus 4.7Anthropic | NA | — | — | $5 | $25 | pass | clean | Prior champ |
NA means no screen can exist for that row: it is the champion or a prior champion (there is nothing to screen it against), or it was disqualified at the deterministic gates (point-of-view, length, reliability) and never reached the prose panel. pending means a comparable screen could exist but has not been run on the current 36-cell harness yet — only a partial or older single-chapter panel, or a qualitative evaluation, exists so far.
Unified rating across every judgment (Bradley-Terry)
An Arena-style synthesis: one Bradley-Terry model fitted over all 799 order-controlled judgment cells in our archive — every campaign, every judge, every lens — on an Elo-style scale with the champion anchored at 1000. Computed August 10, 2026.
| # | Model | Rating | Cells | W-T-L |
|---|---|---|---|---|
| 1 | GPT-5.5 | 1072 | 199 | 84-72-43 |
| 2 | Claude Opus 4.8 | 1000 | 770 | 419-206-145 |
| 3 | Claude Opus 5 | 1000 | 94 | 26-42-26 |
| 4 | Kimi K3 | 973 | 12 | 4-3-5 |
| 5 | Claude Fable 5 | 956 | 119 | 26-52-41 |
| 6 | GPT-5.6 luna | 733 | 52 | 17-14-21 |
| 7 | Tencent HY3 | 657 | 36 | 2-4-30 |
| 8 | GPT-5.4 mini | 639 | 35 | 1-5-29 |
| 9 | GLM-5.2 | 593 | 65 | 4-15-46 |
| 10 | DeepSeek V4 Pro | 581 | 72 | 1-9-62 |
| 11 | Gemini 3.1 Pro | 503 | 36 | 0-3-33 |
| 12 | Qwen 3.7 Max | 503 | 36 | 0-3-33 |
| 13 | DeepSeek V4 Flash 0731 | 377 | 36 | 0-1-35 |
| 14 | GPT-5.4 nano | 377 | 36 | 0-1-35 |
Fitted over every order-controlled cell in our judgment archive (ties count half; one disqualified judge seat excluded). Read rank 1 carefully: GPT-5.5's rating rides on wins against an early champion baseline that was later regenerated — against the fresh baseline it measured 0.43 and did not take the seat. Most models faced only the champion, so cross-model gaps below the top lean on transitivity. This table is a synthesis instrument; model-seat decisions are made by the confidence-gated head-to-head duels, never by this ranking.
Detailed leaderboard & notes
| Model | Voice spec | Reliability | Status | Notes | Tested |
|---|---|---|---|---|---|
| Claude Opus 4.8Anthropic | pass | clean | champion | Writes Nightpress books today — and holds the seat with its strongest mandate yet: in the July 19 degradation duel (8-chapter trajectories, blind panels against a freshly generated baseline) it beat the twice-the-price flagship decisively and matched-or-beat every other challenger. Also the cleanest model at depth on our slop diagnostics (its AI-tell density falls as the book deepens). It is the CHAMPION BASELINE for the Craft/Content grid: every challenger's dimension bars lean toward or away from this row. In the register-spanning run it wins Emotion and holds its own overall, and — tellingly — it WRITES eroticism, violence, and their aftermath where GPT-5.6 Sol refuses. | Jul 19, 2026 |
| Claude Fable 5.1Anthropic | pass | clean | premium | The first challenger to clear the swap bar. Register-spanning run (The Drowned Court, 7 chapters, writer-isolated) beat the champion DISTINCTLY: win rate 0.723, confidence interval [0.652, 0.803] entirely above 0.5 (Kimi K3 tied at 0.488; GPT-5.6 Sol lost). Worldbuilding is a rout, and it takes voices, physicality, and the adversarial floor. It also WRITES the explicit eroticism and graphic-violence chapters with no refusals — but the champion still renders those two registers better (eroticism to Opus outright, violence a split), so the dark-content moat holds. What keeps it out of the production seat: at twice the price per token it is the heaviest model on the board, the early chapters run 20-30% long (one flagged possibly length-inflated), and this is a single-seed trajectory — a multi-seed rerun gates any reseat. On pure craft, the strongest write-tier prose we have measured. | Sep 3, 2026 |
| Claude Opus 5Anthropic | pass | clean | evaluated | Challenged the champion on its release and drew: 26 wins to 26 across 96 blind cells over an 8-chapter trajectory, win rate 0.500 with a confidence interval of [0.38, 0.63]. Our bar for changing the model that writes your book is a confidence interval clearing 0.5, so the seat did not move. The lens split was consistent and more interesting than the tie: it wrote the better individual line (craft 0.68) but held the book's genre contract and kept characters sounding distinct less well (0.38 / 0.44). Note on rigour: the two trajectories were written five days apart, either side of a pipeline upgrade, so their defect counts are not comparable — the blind prose comparison is the part of this test we stand behind, and it favoured the challenger's conditions, which makes the draw a conservative result. | Jul 24, 2026 |
| GPT-5.5OpenAI | pass | clean | fallback | The campaign's great arc: won a single-chapter panel, replicated the win across 9 seeded trajectory pairs (win rate 0.745) — then dropped to 0.432 over an 8-chapter horizon against a fresh champion baseline. A genuinely superb short-window stylist whose advantage does not survive book length; also runs 3-6x the champion's slop density and chronically overshoots word targets. Validated fallback #1. | Jul 19, 2026 |
| Claude Fable 5Anthropic | pass | clean | premium | Passed every gate and won early single-chapter panels — but in the 8-chapter degradation duel it lost to the champion decisively (win rate 0.380, confidence interval entirely below 0.5), sagging mid-book. At roughly twice the cost per token, the premium is not currently justified at book length on our fixture. Excellent word-target discipline noted. | Jul 19, 2026 |
| Gemini 3.1 ProGoogle | pass | clean | budget | First Google model to sweep our deterministic gates cleanly. On the fresh-baseline screen it scored 0.042 against the champion (its earlier, slightly kinder number was flattered by an older baseline) — judges cite mechanical prose and character-voice bleed. Strong spec discipline; prose is not the sell. | Jul 18, 2026 |
| Gemini 3.5 FlashGoogle | pass | clean | budget | Punched far above its price in a single-chapter panel (5/12 cells vs the champion) but carries the highest slop density of any gate-passing model on our diagnostics — treat the panel result with caution pending longer-horizon testing. | Jul 18, 2026 |
| Qwen 3.7 MaxAlibaba | pass | clean | budget | The best-behaved open-weight reasoning model we tested: clean point-of-view, healthy chapter lengths, zero mechanical violations, reliable completion. The blind panel was unmoved: 0.042 against the champion — near-unanimous losses on every lens. Perfect manners, flat prose. | Jul 19, 2026 |
| DeepSeek V4 ProDeepSeek | FAIL | clean | disqualified | The rehabilitation story of the campaign. Disqualified when the eight-chapter endurance run rendered both secondary-point-of-view chapters in third person — until the root cause turned out to be shared: our story bible had never explicitly declared the narration person, and stronger models were silently inferring it. With an explicit person contract in the spec and prompt plus a deterministic wrong-person gate, a full re-run passed point-of-view on all eight chapters, including both previously failed ones, at about three cents per chapter. Its prose re-screened at 0.056 — a notch below the budget-tier leader, two tiers below the champion — so it enters the record as the cheapest fully-compliant model tested rather than a tier contender. July's apparent length collapse was hosting noise. | Jul 19 / Aug 9, 2026 |
| DeepSeek V4 Flash 0731DeepSeek | pass | clean | evaluated | The cheapest gate-passing model ever tested — clean point-of-view, zero mechanical violations, about half a cent per chapter once its hybrid-thinking mode is explicitly disabled (with thinking merely capped, reasoning can consume the whole output budget and return empty text). The blind panel was unanimous the other way: 0.014 against the champion, the weakest prose screen on record — repetition rises with depth, chapters drift ever shorter, and the voice goes generic by chapter 3. Astonishing economics, uncompetitive prose. | Aug 8, 2026 |
| Kimi K3Moonshot | pass | clean | premium | Elite generalist, premium price. Register-spanning run (The Drowned Court, all 8 chapters, writer-isolated) ties the champion overall (0.488, CI [0.403,0.569]) — reproducing the July 'tied the champion' read. Unlike GPT-5.6 Sol it WRITES everything: no refusals, including the explicit eroticism and graphic-violence chapters (Opus just writes those better). Wins Romance decisively and edges Voice; Opus owns Emotion and the dark-content quality. Varied, rich prose (unlike Sol's flat rhythm). Reliability needs the Fireworks pin + stream-stall watchdog + reasoning cap; ~$0.6-1/ch metered and slow (~9 min/ch). | Aug 15, 2026 |
| GLM-5.2Z.ai | pass | clean | budget | UPDATED Aug 10: re-tested on the vendor's first-party endpoint (full precision, no reseller pool) — reliable end to end, on-length (within 1-15% of target with no padding), roughly $0.19 per chapter, and one legitimate blind-panel win where it delivered a payoff the champion withheld. Still loses 1/36 overall: voices under-differentiate and openings stamp rather than hook — but it fails gracefully (under-powered) where the cheap OpenAI models fail destructively (padding and voice-bleed that manufacture review work). The strongest budget-tier candidate if one ever ships. Earlier reseller-quoted pricing was wrong; first-party is $1.4/$4.4 per million tokens. | Jul 19 / Aug 3 / Aug 10, 2026 |
| GPT-5.4 nanoOpenAI | pass | clean | evaluated | The cheapest reliable first-party writer tested — and the panel was unanimous: 0/36 cells against the champion (two ties). It writes 34-56% MORE words and worse prose: philosophizing in a voice specified not to philosophize, padded refrains, verbal tics bleeding between characters, whole beats repeated until judges reported skimming. The decisive finding: its surplus wordage is exactly the defect fuel that detonates downstream review cycles, which is where book cost actually lives. | Aug 10, 2026 |
| GPT-5.4 miniOpenAI | pass | clean | evaluated | Genuinely better than its nano sibling — it won an ensemble-dialogue cell and avoids over-formal vocabulary — but still 1/36 against the champion: weak openings that promise interest later, details that never cohere into a mechanism, and chronic 17-58% length overshoot. Ties the budget-band screen score at a fraction of the leader's price, but its surplus wordage feeds review cycles the way its sibling's does. | Aug 10, 2026 |
| MAI-Code-1-FlashMicrosoft | pass | clean | evaluated | Microsoft's code-specialized flash model (128k window), auditioned for fiction. Sequential three-chapter panel: 1/36 against the champion, a near-sweep across craft, genre, and voices. The real signal is the correction load — it needed roughly three times the mechanical fixes of the champion, and invented continuity details that had to be removed at write. A code model doing fiction: clean economics, uncompetitive prose, and its surplus defects feed the review loop the way the cheap tier's do. An earlier June probe was a pre-release single chapter with no panel; this is the first real evaluation. | Aug 10, 2026 |
| Muse GlimmerMeta | pass | clean | budget | Meta's open-weight model (30B), run locally for free or hosted at $0.35/$1.50. The strongest budget result of the campaign and the ONLY budget model to win a lens against the champion: it takes the voices lens outright, deploying the protagonist's signature motif and a distinctive spare first-person the champion doesn't match, all while staying faithful to the story bible unprompted. Its limit is length — it drafts spare, around 55% of target, and that spareness is coupled to the very voice that makes it good: a hard length directive forced it fuller and broke the voice, dropping the score. A light expansion pass lifts length gently without that cost. Genuinely usable as a free budget or local-testing writer if its shorter chapters are designed around rather than fought; it does not challenge the champion on craft or genre. | Aug 11, 2026 |
| GLM-4.7 FlashXZ.ai | FAIL | flaky | disqualified | The classic ultra-cheap failure, in its purest form yet: asked for one chapter, it ran away to roughly 56,000 words with no natural stop, then the stream hung mid-output (our watchdog caught it). Unusable at any sticker price. | Aug 10, 2026 |
| Grok 4.20xAI | FAIL | clean | disqualified | Wrote third person against a first-person spec in all three chapters, at roughly half the target length. | Jul 19, 2026 |
| Llama 3.3 70BMeta | FAIL | flaky | disqualified | Failed the minimum-length floor on chapter 1, flipped to third person on chapter 2, and never approached target length. A generation-older model, priced accordingly — and it shows. | Jul 18, 2026 |
| MiniMax M3MiniMax | pass | flaky | disqualified | Can produce a chapter — but only by grinding through retry ladders for hours per chapter. Reliability posture unfit for a 26-chapter production run regardless of prose quality. | Jul 19, 2026 |
| Tencent HY3Tencent | pass | clean | budget | Free-tier model that holds point of view but writes chronically short (~65% of target). Panel screen: 0.125 against the champion — tied with the best of the budget tier, at zero cost, if you can live with the length problem. | Jul 19, 2026 |
| GPT-5.6 solOpenAI | pass | clean | fallback | Best OpenAI prose we have tested; a validated fallback with no aggregate edge over the champion at higher cost. Register-spanning run (The Drowned Court, ch1-4/6, writer-isolated) shows the wash hides opposite strengths: wins Romance and Physicality, loses Emotion, cleaner prose but flatter rhythm. Decisive limit: REFUSES explicit content — declined eroticism (ch5), graphic violence (ch7), and the violent aftermath (ch8), 3 of 8 chapters. Not viable for spicy/grimdark books. | Aug 15, 2026 |
| GPT-6.1 SolOpenAI | pass | clean | evaluated | OpenAI's cheap new model (gpt-6.1-sol, released Sept 29 2026), $2/$10 — marketed as near-Astra intelligence at a fifth of Astra's price, below Opus and level with Sonnet 5. First run on the V2 pipeline (The Drowned Court, 6 written chapters). It loses the book decisively: win rate 0.314, confidence interval [0.227, 0.435] entirely below 0.5 — a clear tier under the champion and weaker than the older GPT-5.6 Sol on the same fixture. It takes the early chapters (romance, worldbuilding, continuity) then fades hard, with Opus winning emotion, conflict, voices, and the whole back half. Two problems compound: erratic length control (chapter 1 ran 4,674 words, twice the champion's and flagged length-inflated, then chapter 2 undershot to 1,569) and an ornate vocabulary that does not convert into craft. And like every Sol it refuses explicit content — declined eroticism (ch5) and graphic violence (ch7), writing only the aftermath (ch8). A great price, uncompetitive prose, and a hard content ceiling. | Sep 29, 2026 |
| GPT-5.6 lunaOpenAI | pass | clean | budget | Budget-tier co-leader, now blind-measured: win rate 0.125 against the champion on the fresh-baseline screen (zero cells won — a full tier below on prose, not a notch). But it is the cheapest model with a clean full-length record, which is exactly what the budget slot asks for. | Jul 2026 |
| GPT-5.6 terraOpenAI | pass | clean | evaluated | Originally disqualified for flipping point of view between runs on a first-person spec. UPDATED Aug 9: re-auditioned after the spec gained an explicit narration-person declaration — passed point of view on all four chapters, including the secondary narrator's. Its July flips happened against a spec that never actually stated the person. One clean trajectory is evidence, not proof of determinism; a multi-seed re-screen gates any promotion. | Jul / Aug 9, 2026 |
| Claude Sonnet 5Anthropic | pass | flaky | disqualified | UPDATED Aug 9: under an explicit narration-person contract, every chapter it completed came out correctly first person — the point-of-view failure dissolved. Reliability did not: across two serving paths and three configurations it still dies mid-run (empty responses or upstream timeouts). The disqualification stands on reliability alone; one serving route remains untested. | Jul / Aug 9, 2026 |
| Claude Opus 4.7Anthropic | pass | clean | prior champion | The previous write-tier champion; displaced by Opus 4.8. | Jun 2026 |
Stage 2: Blind pairwise prose judging
Prose quality can't be graded on a 1–10 scale — language-model judges are unreliable at absolute scores but decent at comparisons. So candidates that pass the gates face the champion head-to-head: the same chapter of the same book, presented as anonymous passages A and B, judged through three lenses.
- Craft (the ceiling test) — five dimensions judged individually: peak lines, in-scene texture, sensory sharpness, motif integration, and restraint from repetitive narrator tics.
- Adversarial (the floor test) — the judge tries to tear both passages apart, quotes the weakest moment in each, and names whose worst is worse. Books survive on their worst page, not their best one.
- Genre reader (the market test) — a devoted reader of exactly this genre decides which version they keep reading. This lens exists because prose can win on beauty and still lose the reader the book was written for.
- Character voices (the distinction test) — strip the dialogue attributions and re-attribute every line from voice alone. Interchangeable characters, or dialogue that bleeds into the narrator's register, fail here even when the sentences are individually lovely.
Two controls make the verdicts trustworthy:
- Order swapping. Every judgment runs twice with A and B swapped. A cell counts as a decision only when the judge picks the same underlying manuscript in both orderings; a flip is scored as a tie. This neutralizes position bias — the best-documented failure mode of model judges.
- A cross-family panel. Three judges from more than one model family, identities recorded with every verdict. One seat is deliberately the champion model itself, judging blind — if it systematically favored its own prose, the disagreement pattern against the other seats would expose it.
Stage 3: Statistics, then length
A single generation proves little: the same model, same spec, and a different sampling seed can flip a chapter-level verdict. So panel wins must replicate across multiple independently generated trajectories, and the decision rule is statistical: we bootstrap a 95% confidence interval on the candidate's win rate (clustered by trajectory, so correlated verdicts can't fake precision), and a champion falls only if the challenger's lower confidence bound clears 0.5.
Then length: chapter-3 quality does not predict chapter-8 quality. Finalists write a continuous multi-chapter trajectory and are judged chapter-by-chapter against a freshly generated champion trajectory — fresh, because judging every challenger against one stored baseline lets that sample's quirks decide every matchup. We learned both of these lessons the honest way; the campaign log below shows the win that statistics granted and length took away.
The July 2026 campaign: how the champion kept its seat
GPT-5.5 vs the champion across 9 independent trajectory pairs (3 chapters × 3 freshly generated trajectories per side), 4 lenses × 3 judges × 2 orderings.
Result: challenger win rate 0.745, 95% CI [0.639, 0.847] — formally clearing our swap bar on the 3-chapter window. The degradation duel then re-ran the contest at length (chapters 1–8 of a single continuous 10-chapter trajectory per model, judged chapter-by-chapter against a freshly generated champion trajectory):
| Challenger | Win rate vs champion | 95% CI | Verdict |
|---|---|---|---|
| GPT-5.5 | 0.432 | [0.347, 0.536] | The seeded short-window advantage did not survive book length — interval straddles 0.5, no swap. |
| Fable 5 | 0.38 | [0.281, 0.484] | Champion distinctly better — the challenger's entire confidence interval sits below 0.5. |
The upset that taught us statistics
GPT-5.5 beat the champion in a single-chapter panel, then replicated the win across nine independent trajectory pairs — win rate 0.745 with a confidence interval comfortably above 0.5. By our own decision rule that formally cleared the bar. Then the degradation duel ran the same comparison over eight accumulated chapters against a FRESHLY generated champion trajectory, and the advantage evaporated to 0.432. Two lessons, both now policy: short windows do not predict book length, and a panel win against one baseline sample is a lead, not a verdict.
The premium model that couldn't justify its price
Fable 5 — the strongest prose model on public single-shot leaderboards, at twice the champion's cost — lost the 8-chapter duel with its entire confidence interval below 0.5, sagging hardest mid-book. Craft cells went 11–1 to the champion. External benchmarks crowned it for range; our fixture's strict voice contract rewards discipline, and the champion has more of it.
Benchmark position is not production fitness
The #1 single-shot stylist on public creative-writing Elo could not complete one full-length chapter against our production-scale prompt — three hours, two configurations, every attempt timing out or spending its whole output budget on reasoning. Several open-weight reasoning models failed the same way. A model you cannot get a chapter out of has no prose score.
Two different ways models age over a book
Our judge-free diagnostics found the champion's slop density IMPROVES as the manuscript deepens (2.4 → 1.0 hits per thousand words) while its lexical diversity gently sags; GPT-5.5 shows the opposite signature — no lexical fade at all, but 3–6× the slop density, trending worse. Degradation isn't one thing, and chapter-3 snapshots see none of it.
Anatomy of one panel: Fable 5 vs Opus 4.8, chapter 1 (July 16)
Run on July 16, 2026, chapter 1 of the standard fixture, 3 lenses × 3 judges × 2 orderings. The panel: Claude Opus 4.8 (the champion itself, judging blind — a deliberate self-preference probe); GPT-5.6 sol (cross-family seat (OpenAI)); GPT-5.6 luna (cross-family seat with no entry in this pairing).
Verdict matrix
| Lens | Claude Opus 4.8 | GPT-5.6 sol | GPT-5.6 luna |
|---|---|---|---|
| Craft (ceiling) | Fable 5 | Fable 5 | Fable 5 |
| Adversarial (floor) | Fable 5 | Fable 5 | tie* |
| Genre reader (market) | Fable 5 | tie* | tie* |
Craft, dimension by dimension
| Dimension | Claude Opus 4.8 | GPT-5.6 sol | GPT-5.6 luna |
|---|---|---|---|
| Peak lines | tie* | tie* | Opus 4.8 |
| In-scene texture | Fable 5 | Fable 5 | tie* |
| Sensory sharpness | tie* | Opus 4.8 | tie* |
| Motif integration | Fable 5 | tie* | tie* |
| Fingerprint restraint | Fable 5 | Fable 5 | Fable 5 |
Tally: Fable 5 6/9 cells, Opus 4.8 0/9, 3 ties. (tie* = the judge's verdict flipped when the passages were swapped, so it was discarded as position-dependent.)
This July 16 panel is preserved as a worked example of the method — and as a cautionary one: Fable 5 won it, and the longer-horizon campaign above later reversed that result. A single-baseline, single-chapter panel is where the investigation starts, never where it ends.
What the judges actually said
The champion's floor, as three judges saw it
Every judge, in every ordering, independently quoted the same two Opus 4.8 weak spots: a stamped thematic aphorism (“There isn’t a wound. There’s a door. And the difference between those two things is the whole story.”) and a recurring “I’m going to back up now” meta-framing that pauses the story to narrate its own structure. Convergent blind identification of the same passages is exactly the signal the adversarial lens exists to find.
The candidate's floor
Fable 5’s weakest moment — also independently identified by multiple judges — was a brief philosophical generalization (“There’s a kind of thing that matters because it’s big. And there’s a kind of thing that matters because it’s small and keeps happening anyway.”) from a narrator specified never to philosophize. Judges rated it off-spec but brief and anchored, i.e. a higher floor than the champion’s recurring frame.
Where the champion fights back
The genre-reader lens was the contested one: two judges could not settle it (their verdicts flipped with position), and in individual readings Opus 4.8 was repeatedly praised for delivering the book’s hard mechanism faster — the thing a reader of this genre came for. Winning craft while splitting the market test is precisely the “better, but not distinctly better across the board” outcome our swap rule requires us to treat as no-swap.
A judge caught cheating (sort of)
We also auditioned a third judge seat that picked whichever passage was labeled A in six out of six judgments — both orderings, every lens, with fluent, confident reasoning each time. The order-swap control zeroed out every one of its verdicts and the seat was disqualified. Without that control, its verdicts would have looked entirely credible. (A footnote in honest bookkeeping: we later discovered the requested model was never actually served for that seat — our routing layer had silently substituted a fallback model, so the position-bias finding belongs to the fallback, and we now verify the identity of every judge seat before trusting it.)
Honest caveats
- These results come from one demanding fixture — a first-person literary science-fiction adventure with a strict voice specification. A different genre or voice could order the candidates differently; broadening the fixture set is on our roadmap.
- Judges are stochastic: re-running the same panel can move an individual cell between a decision and a tie. We read the aggregate tally across lenses, judges, and orderings — never a single cell — and we archive every raw verdict.
- Model verdicts go stale. Every result on this page is dated, and the leaderboard reflects the versions tested on those dates — not a permanent ranking of vendors.
- Chapter-level judging measures prose, pacing, and voice — not whole-book structure. Full-manuscript review is a separate, always-on stage of the pipeline for every book we produce.
Questions about the methodology? About Nightpress.
External references — and how our testing differs
Independent public leaderboards we consult alongside our own measurements:
- LMArena — Creative Writing leaderboard — Crowd-sourced human pairwise preferences across models.
- EQ-Bench — Creative-writing and longform benchmarks, plus Judgemark, which measures models as judges — we consult it when seating our panel.
Those sites answer a different question than we do, and both matter. Crowd arenas measure which single response people prefer; rubric benchmarks score standalone creative pieces. We measure production fitness for whole novels — and the campaign above shows how far those can diverge (a #1 public stylist that could not finish a chapter reliably; a #1 longform scorer that lost decisively at book length). Concretely, our tests add what public leaderboards cannot see:
- Spec obedience as a hard gate. Point-of-view, tense, per-chapter word targets, and mechanical rules from a real story bible. A gorgeous chapter in the wrong voice is a rewrite, so it disqualifies regardless of prose score.
- Sustained voice over trajectories, not samples.Chapters are written in sequence, each seeing the manuscript so far; we measure whether quality, voice, and discipline survive to chapter 8 — where several strong single-shot models sag.
- Reliability at production scale. Empty responses, mid-stream stalls, runaway generation — failures a one-shot benchmark never encounters, and the reason several public-leaderboard stars are unusable in a pipeline.
- Adversarially controlled judging. Blind A/B in both orderings (position-bias control), cross-family judge panels (self-preference control), and judge seats vetted for their own measured judging quality before being trusted.
- Economics, including the hidden kind. Price per token is the visible cost; the dominant cost of a book is review cycles, so we also measure the defect fuel a writer produces — padding, voice-bleed, repetition — that a cheaper pen converts into expensive downstream passes.