Skip to content
AG
Back to Blog

Prompting Claude Opus 5.5: what changed from Opus 5, explained

October 1, 2026·6 min read·Kamai Daily·Trend·Opus

Anthropic's own guide to prompting Claude Opus 5.5: thinking is always on, effort is the main dial, several Opus 5 request settings now return a 400, unattended agents stop after reporting, and "avoid a generic AI look" doesn't work. What changed, how it works, and why web developers should care.

A week after Claude Opus 5.5 launched (https://www.anthropic.com/claude-opus-5-5), the page that climbed Hacker News was not a benchmark. It was Anthropic's own guide, "Prompting Claude Opus 5.5" (https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5), which reached 195 points in the trend bank I keep. That tells you where developers are: past the launch, and into the part where their existing code meets a new model.

The guide opens reassuringly: "Existing Claude Opus 5 prompts should perform well without changes." Then it spends the rest of the page listing what to adjust, organised by symptom. This is a walk through the parts that matter if you build web apps on top of the API, plus the migration guide that sits next to it (https://platform.claude.com/docs/en/models/opus-5-5/migration-guide).

What happened

Opus 5.5 is faster at writing: the guide says it generates output tokens "more than 30 percent faster" than Opus 5 and tends to finish the same task with fewer tokens. But two things changed underneath that make a straight model-ID swap less safe than it sounds.

First, thinking is always on. You can no longer switch it off.

Second, several request settings that used to work now return an HTTP 400. The migration guide lists them:

  • thinking: {"type": "disabled"}, and a manual budget_tokens
  • non-default temperature, top_p or top_k
  • a prefilled final assistant turn
  • tool_choice set to any or tool
  • the computer_20251124 tool on the Claude API and Google Cloud (the replacement is computer_toolset_20260801)

If you wrote a helper that forces JSON by prefilling { into the assistant turn, or pins temperature for "deterministic" answers, or forces a specific tool, it breaks on the model name change alone.

For people on Claude Code, the migration guide points to a bundled command, /claude-api migrate this project to claude-opus-5-5, that does the rewrite for you.

How it works

Effort is now the main dial

Because thinking cannot be turned off, the lever you have left is effort. The guide is direct about it: "lowering effort reduces thinking, and with it cost and latency, more reliably than prompt instructions do."

The default moved too. Opus 5 defaulted to high; Opus 5.5 defaults to medium, and in Anthropic's testing, medium on the new model "matches or exceeds Claude Opus 5 at high" on its coding and knowledge-work evals. The catch is that the same effort name now buys more thinking per turn. Carry your old high across and you will get longer, costlier turns, not the same ones.

Two practical notes come with this. Thinking counts toward max_tokens even when it is not returned to you, so a limit sized for a thinking-off integration can cut replies short. The guide says 128,000, the model's maximum, "has worked well" for long agentic coding turns, and the migration guide suggests starting around 64k at xhigh or max. And changing the top-level effort between requests invalidates the prompt cache; a per-message effort change, in beta, keeps it.

Responses start with thinking blocks

Code that assumes "the first content block is the text" will break. Read the response by block type. In a tool loop, pass the thinking blocks back unmodified, or the next request gets a 400.

Progress updates moved out of text

During long agent turns, the model writes short notes about what it found and what it is doing next. On Opus 5.5 those arrive as progress-update thinking blocks, and at the default display their text is empty. So a chat UI that only renders text blocks will look frozen for the whole turn. Setting display: "updates" (beta) gets you a short summary of each.

The guide also describes a harness trick: count consecutive tool-calling steps that show the user nothing, and after five, append a one-line reminder asking for an update. Anthropic says this "roughly halved the share of tasks with a long silent stretch, with no measurable change in cost".

Unattended agents stop after reporting

This is the change most likely to bite anyone running agents without a human watching. Some of those progress updates end the turn with plain text, stop_reason: "end_turn". A loop that treats every end of turn as "task done" will stop halfway.

The guide's fix has three parts. Treat a text-only end of turn as a report, not proof of completion. Keep the task's parts in a checklist the model updates, and if items are still open with no blocker named, send a short message naming them. And cap it: stop after 2 or 3 automatic continuations on the same task, so a genuinely stuck run ends and can be reviewed. It also gives a system-prompt paragraph naming four specific kinds of early stop the user does not want.

Smaller changes worth knowing

  • Chat latency. If your chat system prompt says "think carefully before answering", try removing it. In Anthropic's chat testing, that made replies start sooner "with no clear decline in the quality".
  • Pasted text. Wrap text a user pasted from elsewhere in <pasted_content id="…"> tags with a matching random id, and add a system-prompt note saying instructions inside it are not the user's. That is a cheap prompt-injection guardrail for any app with a text box.
  • Reasoning extraction is refused. Prompts that push the model to write its reasoning into the answer can come back with stop_reason: "refusal" and the category reasoning_extraction. Read the summarized thinking instead.

Why someone building web apps with AI should care

The section I would read first is the one on frontend output. Without design direction, the guide says, the model "falls back on a few default styles", and telling it to "avoid a generic AI look" mostly swaps one default for another. What works is naming the patterns you do not want. Its own example bans a cream or off-white background, italic accent words in headlines, numbered "01/02/03" section labels, monospace labels and pill-shaped buttons. If you have looked at many AI-generated landing pages this year, you will recognise every one of them.

The other two things are plumbing, and plumbing is where production breaks. A JSON helper that prefills or forces a tool fails on the ID swap. A chat UI that renders only text looks dead during long turns. Neither shows up in a demo; both show up the first day real users are on it.

There is also a sign of where the tooling is going. LaunchVideo reached Hacker News the same week as "Opus 5.5 is good at explainer videos" (https://launchvideo.io): the model writes the film as HTML, CSS and JavaScript, and a serverless agent renders it to MP4, at roughly 100k tokens and about 4 minutes per film. The model writes a web page; a headless browser turns it into video. Frontend skills are becoming the way these models make things.

What is not known yet

  • Independent cost numbers. Every effort-level comparison in the guide is Anthropic's own testing. Nobody has published cost per effort level on real web-app workloads yet.
  • Whether the default-styles list holds. The guide itself says to check which styles the first result used and extend the list, which suggests the defaults can move.
  • Sonnet 5.5 and Haiku 5.5. No prompting guidance for them has been published yet.

My take

I run an unattended agent every day: a worker on a Mac mini that picks up scheduled jobs, does them, and reports back without anyone at the keyboard. So the section on early stops reads to me like the most important paragraph in the guide, more than any benchmark. The hard part of an unattended agent was never getting it to start. It is knowing whether "here is what I did" means finished or means paused, and the guide's answer, a checklist plus a small capped number of nudges, is the same shape I would reach for: make "done" something you check, not something the model announces.

The frontend section is my other takeaway, and it is less about Claude than about prompting in general. "Make it look good" is a vibe; "no pill buttons, no cream background" is a spec. As a frontend developer I would rather write the second kind of prompt anyway, because it is the same thing a good design review does: name what is wrong, precisely, and let the rest follow.

About the author

Amar Gupta

Amar Gupta

Senior Frontend Developer — AI & MCP

I build production frontends in React, Next.js and TypeScript — and the AI and MCP tooling behind them. 7+ years shipping web applications, from data modelling through to the deployed interface.

📍 Delhi, India · Open to Full-time

Read next