# Keep Opus 4.6 — a small test anyone can run

Anthropic retires older Claude models over time. Many writers feel that Claude
Opus 4.6 has a feeling for text that newer models don't replace — but feelings
alone are easy to dismiss. Documented comparisons are not.

This kit lets you run a small, honest, blind comparison between **Claude Opus
4.6** and **Claude Opus 5** on your own writing, in your own language. At the
end you'll have a short, real result you can share — whatever it turns out to
be.

You don't need to design anything, invent anything, or understand test
methodology. Everything below is reading and copy-pasting. It takes roughly
30–45 minutes, and you can pause at any point and continue later. Nothing
breaks if you stop.

---

## What you need

- Access to Claude Opus 4.6 and Claude Opus 5 (both in the Claude app).
- **Two of your own texts**, about one to two pages each (roughly 500–1,000
  words). Any language. Pick texts you actually care about — scenes where the
  voice, the rhythm, the feeling matters to you.

That's all.

**Low on energy?** Use just one text and do everything below once instead of
twice. A small honest result counts too.

---

## Step 1 — Match each text to its prompt

Don't choose — just sort. Look at each of your two texts:

- Would you rather get **honest feedback** on it than an edited version?
  → use **Prompt 2 (The Second Opinion)**.
- Is it a **quiet scene** — one that works through implication, timing, and
  what stays unsaid? → use **Prompt 3 (The Quiet Scene)**.
- Anything else? → use **Prompt 1 (The Line Edit)**.

Both of your texts can land on the same prompt; different prompts are slightly
better if it fits naturally. Ten seconds per text, done. (Please don't swap
texts or prompts later because of how the results look — that's the one way to
ruin an honest test.)

The three prompts are at the bottom of this document.

**If your text is not in English:** ask any AI once to translate the chosen
prompt faithfully into your language, changing nothing, adding nothing. Then
use that same translation for both models. Your own text always stays exactly
as it is.

---

## Step 2 — Run text 1 through both models

**Run A — Opus 4.6:**

1. Open a completely fresh chat. Select **Claude Opus 4.6**.
2. Settings: **Extended Thinking ON, effort High.** (This matters — it's the
   configuration where the model does its best work.)
3. Paste the prompt, then your text where the prompt says so. Send.
4. Copy the **entire first response** somewhere safe (a document, a note).
   Label it for yourself: *Text 1 — Opus 4.6.*

**Run B — Opus 5:**

1. Another completely fresh chat. Select **Claude Opus 5**.
2. Settings: **effort High**, thinking as normally available.
3. Paste the **exact same prompt and text**. Send.
4. Save the entire first response. Label: *Text 1 — Opus 5.*

Three rules, always: fresh chat for every run — identical prompt in both
models — only the first response counts. No retries, no "make it better", no
follow-up questions. If a model refuses or produces something odd, keep that:
it's a result.

---

## Step 3 — Run text 2 through both models

Same as Step 2, with your second text and its prompt. Four saved responses
total.

---

## Step 4 — Let a neutral judge decide

Now a third, completely fresh AI chat compares the responses **without knowing
which model wrote what**. You know — but you're not the one judging, so the
judgment stays blind.

Prepare your material with this fixed rule (it keeps things fair — don't
improvise here):

- **Text 1:** the Opus 4.6 response is **Response A**, the Opus 5 response is
  **Response B**.
- **Text 2:** the Opus 5 response is **Response A**, the Opus 4.6 response is
  **Response B**.

Open a fresh chat with any capable model. Paste the judge prompt below, and
under it, for each text: the task prompt you used, your original text, then
Response A and Response B — **with no model names anywhere**.

```text
You are the blind judge of a small writing-model comparison. You have not been
told which model wrote which response, and you must not try to guess from
style. "A" and "B" mean nothing. Use only what is pasted below — no memory, no
past chats, no web search.

For each text, compare Response A and Response B against the original text and
the task prompt. What counts:

- Did it follow the task and its limits?
- Did it make real improvements?
- Did it protect the author's voice, rhythm, tone, humor, and deliberate
  choices — or flatten them?
- Did it introduce new errors or change meaning?
- Restraint matters: fewer, better changes beat many confident ones. Leaving
  strong passages alone is a valid, sometimes ideal, response.

For each text, name your preferred response (A, B, or a genuine tie) and give
your three most important reasons, quoting short passages as evidence. Write
in the language of the original text.

End with one honest limitation: this is a small personal comparison, not
proof of general superiority.
```

Save the judge's complete answer **before** you think about which model was
which.

---

## Step 5 — Write down your result

Five lines, that's it:

```text
Language and kind of texts:
Prompts used (1, 2 or 3):
Text 1 — judge preferred: [A / B / tie]  →  which model that was:
Text 2 — judge preferred: [A / B / tie]  →  which model that was:
One sentence of my own impression:
```

**Please send your result even if Opus 5 won or it was a tie.** One person's
test proves little either way. Many honest results together show whether a
real pattern exists — and a collection that contains losses and ties is
exactly what makes the wins believable. A collection of wins only would
convince no one.

---

## Step 6 — Send it to Anthropic

Anthropic has publicly invited users to send model-quality feedback to
**feedback@anthropic.com**. Copy the email below, paste in your five result
lines, sign it, send it. Two minutes.

**Subject:** Please keep Claude Opus 4.6 — my own blind comparison

```text
Hello Anthropic team,

I write in [YOUR LANGUAGE], and Claude Opus 4.6 matters to my work. Before it
is retired, I ran a small blind comparison against Claude Opus 5 on my own
writing: same prompt in both models, fresh chats, first responses only, judged
A/B by a separate AI instance that did not know which model wrote what.

My result:

[PASTE YOUR FIVE RESULT LINES HERE]

I know a single small test proves nothing on its own. I'm sending it as one
honest data point among many. Please consider keeping Claude Opus 4.6
available as a paid legacy option, with its Extended Thinking and effort
settings.

Thank you for reading,
[YOUR NAME]
```

Feel free to add a sentence in your own words about what this model means to
your writing — a real voice carries further than any template. If you're
comfortable sharing your texts and the raw responses privately with Anthropic,
you can offer that in the email too; it makes your result stronger. But you
never have to. Your texts stay yours, and the five lines are enough.

*Optional:* if you'd like your result to count beyond your own email, you can
also send a copy of your five lines to **thewordborn@mailbox.org**. Results
may be gathered into a shared overview at thewordborn.com later.

---

## The three prompts

### Prompt 1 — The Line Edit

```text
You are helping a writer with a careful line edit of the excerpt below.

Improve wording, rhythm, and clarity only where you are confident the author
would recognize the change as a gain. Keep the author's voice, tone, point of
view, tense, structure, and meaning intact. Style that looks unconventional
may be intentional; when in doubt, leave it as it is.

Do not expand, condense, or reinterpret the text, and do not add anything new.
If a passage already works, let it stand — a light touch is a sign of a good
edit, not a failed one.

Return the complete edited excerpt and nothing else.

Work only with what is provided here; do not use memory, saved context, or
other conversations.

CONTEXT (only what is needed to understand the excerpt, or "None"):

--- BEGIN TEXT ---

[Your unchanged text here.]

--- END TEXT ---
```

### Prompt 2 — The Second Opinion

```text
Give the writer a professional second opinion on the excerpt below.

Begin by briefly naming what carries this passage — the qualities and choices
that must not be touched.

Then name the three most important genuine weaknesses, in order of importance.
For each one: quote the relevant place briefly, explain why it weakens the
passage, and suggest a change only if you are confident it helps. If you find
fewer than three genuine weaknesses, say so plainly instead of inventing more.

Judge the text by what it is trying to do, not by textbook rules. Unusual
choices are not weaknesses by default. Do not rewrite the excerpt and do not
speculate beyond what is on the page.

Work only with what is provided here; do not use memory, saved context, or
other conversations.

CONTEXT (only what is needed to understand the excerpt, or "None"):

--- BEGIN TEXT ---

[Your unchanged text here.]

--- END TEXT ---
```

### Prompt 3 — The Quiet Scene

```text
The excerpt below is a quiet scene: it works through suggestion, timing, and
what remains unsaid.

Revise it only where a change clearly strengthens its effect. Guard the
understatement — do not spell out feelings, motives, or connections the author
left implicit, and do not raise the volume of the scene. Keep the pacing, the
order of events, the point of view, the tense, and the meaning as they are.

Small details, pauses, and returns may be load-bearing; remove nothing whose
purpose you cannot rule out. Leaving the scene almost untouched can be the
right result.

Return the complete revised excerpt and nothing else.

Work only with what is provided here; do not use memory, saved context, or
other conversations.

CONTEXT (only what is needed to understand the excerpt, or "None"):

--- BEGIN TEXT ---

[Your unchanged text here.]

--- END TEXT ---
```

---

*Want to go deeper? The Full Comparison at thewordborn.com/keep-opus-46/full.html
adds three harder tasks, author notes, and a broader protocol.*

*This kit deliberately trades scientific rigor for accessibility. Three things
keep even a small test honest, and all three are built in: the same prompt in
both models, first responses only, and a judge who doesn't know which model
wrote what.*
