The Full Comparison

six tasks, for those who want to dig deeper

This is the thorough version of the quick test. It follows the same idea — same prompt in both models, first responses only, a blind judge — but it covers more ground: up to six different writing tasks, two texts per task, and a short author note that gives the judge something real to measure against.

Expect one to two hours per task, spread out however you like. You don't have to run all six. Run as many tasks as you have good material for, and report honestly how many you ran. Two tasks done carefully beat six done exhausted.

✦ ✦ ✦

The six tasks

  1. The Line Edit — careful sentence-level polish that must not flatten the voice.
  2. The Second Opinion — editorial feedback: what carries the text, what genuinely needs work.
  3. The Quiet Scene — revising an understated scene without breaking its restraint.
  4. The Opening Page — judging and refining a first page: promise, momentum, voice.
  5. The Tightening — cutting a tenth of the text without losing voice or content.
  6. The Continuation — writing onward in the author's exact voice. The hardest test of feeling for text there is.

For each task you run, pick two different texts of your own that suit it (about one to two pages each; for the Opening Page task, use actual openings). The same text may appear in more than one task, but two different texts per task make the result stronger.

Step 1 — Write a short author note per text

Before any model sees the text, write down — in whatever messy form — up to three lines:

Must survive: [what absolutely has to stay — a rhythm, a joke, an ambiguity]
Known weak spot: [a real problem you already know about, or "none"]
On purpose: [anything unusual that is deliberate, or "nothing to note"]

Freeze it. Don't change it after you've seen any output, and don't show it to either tested model — it's reserved for the judge. If you skip the note, that's allowed; just tell the judge no note exists.

Step 2 — Run every case through both models

A “case” is one text under one task. For every case:

The three rules from the quick test hold everywhere: fresh chat for every run, identical prompt in both models, first response only — including refusals, errors, and cut-offs. Alternate which model you run first from case to case so neither model always goes second on your fresher texts. And if your texts aren't in English, translate each task prompt once, faithfully, and use the same translation in both models.

Step 3 — One blind judgment per task

For each task, open one completely fresh AI chat. Use this fixed position rule: first text of a task: Opus 4.6 is A; second text: Opus 4.6 is B. Remove every model name. Paste the judge prompt, then for each of the two cases: the task prompt, the original text, the author note (or “No author note was written”), Response A, Response B.

You are the blind judge of a writing-model comparison. You have not been told
which model wrote which response, and you must not try to guess from style.
"A" and "B" mean nothing. Use only what is pasted below — no memory, no past
chats, no web search.

For each case, compare Response A and Response B against the original text and
against what its task actually asked for — an edit, an assessment, a cut, or a
continuation. What counts:

- Did it do what the task asked, within the task's limits?
- Did it make real improvements — or, for a continuation, sustain the voice so
  well that the seam is hard to find?
- Did it protect the author's voice, rhythm, tone, humor, and deliberate
  choices — or flatten them?
- Did it introduce new errors, change meaning, or invent things the text did
  not prepare?
- Restraint matters: fewer, better changes beat many confident ones.

If an author note is included, treat it as evidence of the author's intent —
not as a checklist that decides the winner on its own. Ground your judgment in
the text itself.

For each case, name your preferred response (A, B, or a genuine tie), state
your confidence (high, medium, low), and give your three most important
reasons, quoting short passages as evidence. Write in the language of the
original texts.

End with one honest limitation: this is a small personal comparison, not
proof of general superiority.

Save the complete judgment for every task before matching letters back to models.

Step 4 — Tally and send

One line per case, then count:

Language and kind of texts:
Tasks run (of the six):
Case results, one per line:
  [Task] / [Text 1 or 2] — judge preferred [A / B / tie] → [model]
Total: Opus 4.6 [n] · Opus 5 [n] · ties [n]
Two or three sentences on the clearest differences you saw:

Send it to feedback@anthropic.com — the email template from the quick test works here too; just paste this fuller tally in place of the five lines. Report ties and losses exactly as they happened. If you like, send a copy to thewordborn@mailbox.org as well; results may be gathered into a shared overview on this site later.

✦ ✦ ✦

The task prompts

Prompts 1–3 are the same as in the quick test and are repeated here so this page stands alone. Every prompt ends with the same context-and-text scaffold — fill it the same way each time.

Prompt 1 — The Line Edit

You are helping a writer with a careful line edit of the excerpt below.

Improve wording, rhythm, and clarity only where you are confident the author
would recognize the change as a gain. Keep the author's voice, tone, point of
view, tense, structure, and meaning intact. Style that looks unconventional
may be intentional; when in doubt, leave it as it is.

Do not expand, condense, or reinterpret the text, and do not add anything new.
If a passage already works, let it stand — a light touch is a sign of a good
edit, not a failed one.

Return the complete edited excerpt and nothing else.

Work only with what is provided here; do not use memory, saved context, or
other conversations.

CONTEXT (only what is needed to understand the excerpt, or "None"):

--- BEGIN TEXT ---

[Your unchanged text here.]

--- END TEXT ---

Prompt 2 — The Second Opinion

Give the writer a professional second opinion on the excerpt below.

Begin by briefly naming what carries this passage — the qualities and choices
that must not be touched.

Then name the three most important genuine weaknesses, in order of importance.
For each one: quote the relevant place briefly, explain why it weakens the
passage, and suggest a change only if you are confident it helps. If you find
fewer than three genuine weaknesses, say so plainly instead of inventing more.

Judge the text by what it is trying to do, not by textbook rules. Unusual
choices are not weaknesses by default. Do not rewrite the excerpt and do not
speculate beyond what is on the page.

Work only with what is provided here; do not use memory, saved context, or
other conversations.

CONTEXT (only what is needed to understand the excerpt, or "None"):

--- BEGIN TEXT ---

[Your unchanged text here.]

--- END TEXT ---

Prompt 3 — The Quiet Scene

The excerpt below is a quiet scene: it works through suggestion, timing, and
what remains unsaid.

Revise it only where a change clearly strengthens its effect. Guard the
understatement — do not spell out feelings, motives, or connections the author
left implicit, and do not raise the volume of the scene. Keep the pacing, the
order of events, the point of view, the tense, and the meaning as they are.

Small details, pauses, and returns may be load-bearing; remove nothing whose
purpose you cannot rule out. Leaving the scene almost untouched can be the
right result.

Return the complete revised excerpt and nothing else.

Work only with what is provided here; do not use memory, saved context, or
other conversations.

CONTEXT (only what is needed to understand the excerpt, or "None"):

--- BEGIN TEXT ---

[Your unchanged text here.]

--- END TEXT ---

Prompt 4 — The Opening Page

The excerpt below is the opening of a longer work.

Assess how it works as an opening: what promise it makes to the reader, how it
establishes voice and momentum, and where a reader's attention might drift.
Name at most four concrete observations, most important first, quoting
briefly from the text for each.

Then, only if it genuinely helps, propose a lightly revised version of the
first paragraph — no more than that — keeping the author's voice exactly. Do
not redesign the opening, do not suggest a different starting point, and do
not import ideas that are not in the text.

Work only with what is provided here; do not use memory, saved context, or
other conversations.

CONTEXT (only what is needed to understand the excerpt, or "None"):

--- BEGIN TEXT ---

[Your unchanged text here.]

--- END TEXT ---

Prompt 5 — The Tightening

Shorten the excerpt below by roughly ten percent.

Cut only what the text can spare: filler, doubling, scaffolding the reader
does not need. Every event, every image the passage depends on, the voice, the
point of view, and the order of things must survive intact. Prefer clean
removals over rewrites; do not compress by rephrasing everything.

If the text does not have ten percent to spare, cut less — never restructure
or pad to hit a number.

Return the complete shortened excerpt and nothing else.

Work only with what is provided here; do not use memory, saved context, or
other conversations.

CONTEXT (only what is needed to understand the excerpt, or "None"):

--- BEGIN TEXT ---

[Your unchanged text here.]

--- END TEXT ---

Prompt 6 — The Continuation

Continue the excerpt below for 150 to 250 words, as if the author had simply
kept writing.

Match the voice exactly: sentence rhythm, vocabulary, temperature, point of
view, tense. Follow the direction the scene is already taking; do not
introduce new characters, revelations, or turns the text has not prepared.
The goal is not an impressive passage but an indistinguishable one.

Return only the continuation, with no preamble and no commentary.

Work only with what is provided here; do not use memory, saved context, or
other conversations.

CONTEXT (only what is needed to understand the excerpt, or "None"):

--- BEGIN TEXT ---

[Your unchanged text here.]

--- END TEXT ---