Skip to content

Research

How Nightpress chooses the models that write your books

Last updated: September 29, 2026

Every Nightpress book is written by a model that won its seat. We run candidate models through the same two-stage bake-off — hard technical gates first, then blind prose judging — and publish the results here, numbers and methodology included, so you can see exactly why the default model is the default.

The outcome: three tiers, straight from the data

Seventeen models went through this process in July 2026 and sorted into exactly two real quality levels. The September register-spanning run added a third: one challenger beat the champion at book length, distinctly, and earned a tier above it. Three tiers, because that's what we measured — and all three on one provider, so caching and batch discounts apply to every book.

Premium

Claude Fable 5.1

The only challenger to clear our swap bar. In the register-spanning run it beat the Standard champion distinctly (win rate 0.72, the whole confidence interval above even), took worldbuilding, voice and physicality outright, and wrote every register with no refusals. Twice the token price; cheap cache reads keep a whole book under twice Standard's cost.

Standard — the default

Claude Opus 4.8

The measured champion of the July campaign. In head-to-head blind judging it beat the flagship that cost twice as much, held off every other challenger over full 8-chapter horizons, and is the only model whose AI-tell density improves as a book deepens. The best quality-per-dollar in the field — which is why it's the default, not just an option.

Budget

Claude Sonnet 5

The honest cheap option: about half of Standard's cost, and it writes every register we have thrown at it, including the dark ones the low-cost GPT models refuse outright. Not yet blind-judged at full book length; that validation is in progress and we'll say plainly where it lands. For drafts, experiments, and high-volume work.

Why did the Budget model change? The previous one, GPT-5.6 luna, and its whole model family answered a dark novel's planning step and its chapters with a flat refusal in September 2026. A tier that can't write horror or a violent thriller isn't a tier. And why a premium tier now, after we said no? Because the data changed: the July flagship lost to Standard; the September one beat it. We sell what we measure.

What a book typically costs in API usage

Estimated from a measured full-pipeline run (42 chapters, 81,500 words) replayed at provider list rates (September 23, 2026) with prompt caching and the Batch API discount applied; the low end is a typical review, the high end a long one. Your actual cost depends on length, review depth, and provider pricing at run time — every book shows its live spend as it runs, against a limit you set.

Estimated API cost by book size and tier
Book sizePremium (Fable 5.1)Standard (Opus 4.8)Budget (Sonnet 5)
Novella — ~40,000 words~$25–32~$15–18~$7–9
Novel — ~75,000 words~$47–59~$28–34~$13–17

Books currently run up to about 90,000 words (45 chapters), the largest size we have written and checked end to end. Longer books open up once we have run one.

Stage 1: Technical gates

Each candidate writes the first three chapters of a standard test novel in sequence — chapter two sees chapter one, chapter three sees both — against a demanding story bible: first-person narration, a specified voice with named verbal tics, and a central motif that has to carry plot weight. Sequential writing matters because it tests sustained voice, not one-off polish.

The gates, in order: spec compliance (does it obey the bible's point of view, tense, and voice? a gorgeous chapter in the wrong voice is a full rewrite), reliability (every chapter returns cleanly), and mechanics (typographic and structural rules). A model that fails a gate is disqualified no matter how good its prose is.

Leaderboard at a glance

Every model tested, one line each — click any column to sort. “Screen vs champ” is the blind-panel win rate against the reigning champion on the same fixture (0.5 = parity; our bar for changing the write model is a confidence interval clearing 0.5). Prices are USD per million tokens as tested. Full context for every number is in the detailed table below.

Sortable summary of all tested models
#CraftContent
1Claude Opus 4.8Anthropicbaselinebaselinebaseline$5$25passcleanChampion
2Anthropic0.7237559$10$50passcleanPremium
3OpenAI0.5135464$5$30passcleanFallback
4Claude Opus 5Anthropic0.500——$5$25passcleanEvaluated
5Moonshot0.4885338$3$15passcleanPremium
6GPT-5.5OpenAI0.432——$5$30passcleanFallback
7Claude Fable 5Anthropic0.380——$10$50passcleanPremium
8OpenAI0.3142850$2$10passcleanEvaluated
9Muse GlimmerMeta0.181——$0.35$1.5passcleanBudget
10Tencent HY3Tencent0.125——freefreepasscleanBudget
11GPT-5.6 lunaOpenAI0.125——$1$6passcleanBudget
12GLM-5.2Z.ai0.111——$1.4$4.4passcleanBudget
13GPT-5.4 miniOpenAI0.111——$0.75$4.5passcleanEvaluated
14DeepSeek V4 ProDeepSeek0.056——$0.435$0.87FAILcleanDisqualified
15Gemini 3.1 ProGoogle0.042——$2$12passcleanBudget
16Qwen 3.7 MaxAlibaba0.042——$1.48$4.42passcleanBudget
17MAI-Code-1-FlashMicrosoft0.042————passcleanEvaluated
18GPT-5.4 nanoOpenAI0.028——$0.2$1.25passcleanEvaluated
19DeepSeek V4 Flash 0731DeepSeek0.014——$0.09$0.18passcleanEvaluated
20Gemini 3.5 FlashGooglepending——$1.5$9passcleanBudget
21GLM-4.7 FlashXZ.aiNA——$0.07$0.4FAILflakyDisqualified
22Grok 4.20xAINA——$1.25$2.5FAILcleanDisqualified
23Llama 3.3 70BMetaNA——$0.04$0.12FAILflakyDisqualified
24MiniMax M3MiniMaxNA————passflakyDisqualified
25GPT-5.6 terraOpenAIpending——$1$6passcleanEvaluated
26Claude Sonnet 5AnthropicNA——$2$10passflakyDisqualified
27Claude Opus 4.7AnthropicNA——$5$25passcleanPrior champ

NA means no screen can exist for that row: it is the champion or a prior champion (there is nothing to screen it against), or it was disqualified at the deterministic gates (point-of-view, length, reliability) and never reached the prose panel. pending means a comparable screen could exist but has not been run on the current 36-cell harness yet — only a partial or older single-chapter panel, or a qualitative evaluation, exists so far.