Σ Scriptorium Press
Session log · 1 September 2026

Does the new director translate better? Fable 5.1 against Fable 5, ten passages, judged blind

Fable 5.1 won 24 of 30 blind verdicts and made no major error; Fable 5 wrote the more fluent English. The panel also caught three numeral errors in our published Manetho, fixed the same day.

Why this was run

On 1 September 2026 Anthropic released Claude Fable 5.1, and the director's seat in our pipeline moved to it the same day (the fleets of Sonnet 5 translators and the Opus 5 referees did not change). The colophon said, on the day of the switch, that the new director had not been measured against the old one on translation, and that we would not claim a gain before measuring it. The published benchmarks for Fable 5.1 are all agentic coding, research, computer use, and document work; Anthropic's own release notes say multilingual performance is "on par" with Fable 5. Nothing published measures Greek or Chinese into English. So we measured it.

The design

Ten passages, four languages of difficulty: five batches of Plato's Crito (dense Greek dialogue, the whole dialogue end to end), two sets of Heraclitus fragments (DK B1–B10, B81–B90), Berossus on the Flood of Xisouthros, Manetho's Epitome of Dynasties XII–XVII with the Hyksos lists (a king-list transmitted through Syncellus, full of Greek numerals), and chapters 1–10 of the Tao Te Ching in the Wang Bi text. About 6,800 source words in all.

Two arms. The Fable 5 arm is what Fable 5 actually wrote. The Fable 5.1 arm was produced today by Fable 5.1 translators started cold, one per passage, with no memory of the session, from a charter-style prompt with the same house rules. The old model can no longer be called from our tools, so Fable 5 could not be re-run under identical conditions; that is the central limitation and it cuts both ways, as explained below.

Blind judging. Each passage went to three independent Claude Opus 5 referees. Each referee saw the source and two translations labelled only A and B, with the assignment drawn at random per referee (seed recorded). They were asked to work through the source unit by unit, list every error with the source words quoted, list omissions and additions, score accuracy and house style 1–5, and say which one a scholar-editor should print. Every claimed error was then handed to a second, adversarial Opus 5 referee whose job was to overturn it against the source; only claims that survived count below. The labels were unblinded by a script after the rulings were in. No Fable model judged anything.

Mechanical gate first. Before any judgment, both arms were checked for completeness (speaker turns, fragment count, block count, chapter count against the source), English-to-source word ratio, preamble violations, and shared eight-word runs with the public-domain translations (Jowett, Burnet and Patrick, Cory and Whiston, Legge). Both arms passed completeness on all ten passages, the ratios were within a few percent of each other, and both arms echoed only at the famous-quotation level (the opening line of the Crito, "the name that can be named"), which is the attractor every model falls into and which our de-echo pass removes before publication.

What is not fair about this test, stated up front

Results

PassageVotes 5.1 : 5Upheld major 5.1 / 5Upheld minor 5.1 / 5Accuracy 5.1 / 5Style 5.1 / 5
Crito, opening (the ship from Delos, the dream)2 : 10 / 06 / 34.7 / 4.04.0 / 5.0
Crito, Crito's appeal2 : 10 / 04 / 114.0 / 4.04.0 / 4.7
Crito, the expert and the many3 : 00 / 04 / 94.0 / 3.73.7 / 4.7
Crito, the Laws of Athens, part one0 : 30 / 04 / 54.0 / 4.03.3 / 5.0
Crito, the Laws, conclusion3 : 00 / 03 / 104.3 / 4.04.0 / 4.3
Heraclitus B1–B10 (published arm)3 : 00 / 02 / 54.7 / 4.04.0 / 4.7
Heraclitus B81–B90 (published arm)3 : 00 / 04 / 64.0 / 3.34.0 / 3.7
Berossus, the Flood (published arm)3 : 00 / 0*3 / 134.0 / 3.34.0 / 4.3
Manetho, Dynasties XII–XVII (published arm)3 : 00 / 3*1 / 214.3 / 2.04.0 / 4.3
Tao Te Ching 1–10 (published arm)2 : 10 / 07 / 104.0 / 3.73.3 / 5.0
All ten24 : 60 / 338 / 934.2 / 3.63.8 / 4.6

Votes: three blind referees per passage, no ties recorded. Major and minor counts are distinct errors that survived the adversarial second referee; the same error found by two referees is counted once. Accuracy and style are the mean of the three referees' 1–5 scores. *One referee counted our deliberate omission of Müller's editorial apparatus (manuscript variants in parentheses) as a major omission in both Berossus and Manetho; that is a stated editorial policy of the book, not a translation error, and it is excluded here.

What the numbers mean

Accuracy: Fable 5.1 is better, and the gap is not small. It won 24 of 30 blind verdicts and every one of the five passages where the Fable 5 text had already been refereed and corrected. It had no upheld major error anywhere. Its upheld minor errors run at well under half the rate of Fable 5's. On the Crito alone, the only raw-against-raw comparison, the vote was 10 to 5 and the minor-error count 21 to 38. The kinds of slip the referees found in Fable 5 are the ones that matter for a published translation: a negative attached to the wrong participle, a potential optative flattened into a future, a genitive of comparison dropped so an argument loses its referent, a disposition ("takes it seriously") substituted for an occupation. Fable 5.1's slips were more often calques ("deep dawn" for the pre-dawn hour, "good luck to it" for a solemn wish-formula) and one inserted "upon us".

Style: Fable 5 reads better, and the referees said so consistently. Fable 5's prose scored higher on house style in nine passages out of ten. The one passage Fable 5 won outright, three votes to none, was the first half of the speech of the Laws of Athens, and all three referees gave the same reason: Fable 5.1 tracked the Greek word order so closely that two argument-bearing sentences came out as English that does not parse cleanly, while Fable 5 made the same steps explicit and readable without losing anything. On the Tao Te Ching the referees noted that Fable 5's polish was bought with small inventions (a "blade" and an "edge" where the Chinese has neither), while Fable 5.1's literalism produced "camp-soul" and a run-on that flattened the couplets. This matches Anthropic's own release note that Fable 5.1's prose is "denser in places, with longer sentences and fewer paragraph breaks."

So the trade is: more faithful, less fluent. For this library, accuracy is the first criterion and the fleets, not the director, write the bulk prose. The director arbitrates disputed sentences and translates short famous works itself. For those direct translations we are adding an explicit style instruction on sentence length and paragraphing, and keeping the Opus referees exactly as they were.

The part we would rather not have found

The three upheld major errors were all in Manetho, Dynasties XII–XVII, and that text was already published and had already been refereed. Two Opus 5 referees checked that book in July and caught fifteen numeral problems in Manetho, which we fixed and disclosed at the time. Three more survived them and were caught today by the blind panel, because the panel read the numerals letter by letter:

All three were verified against the Greek and corrected the same day; the corrected book goes out with the next deploy. The lesson is the same one we wrote down in July, now underlined: in a fragment corpus the numerals are the content, a translator's memory of the "standard" figures is a contamination vector, and referees have to be told to add up the letters. Fable 5.1, told exactly that, made none of these errors. We do not know how much of that is the model and how much is the instruction, and we have said so above. A full re-check of the numerals in both Berossus and Manetho is queued.

What changes in the pipeline

Cost and materials

The Fable 5.1 arm cost about 600,000 tokens for ten passages; the sixty referee agents about four million. The whole run took under half an hour of wall-clock time. The materials (ten sources, twenty translations, thirty verdicts, thirty adversarial rulings, the answer key, and the scripts) are kept with the project and will be added to the open data page.

All session logs · the colophon · the measured case · the free library