Blog · 29 September 2026 · 7 min read

Is Claude Sonnet 5.5 safe? The system card and your data

Claude Sonnet 5.5 shipped on 28 September 2026. What Anthropic’s system card says about its risks, and what your plan decides about your data.

New model. Same data rules. Check the toggle.

Short answer: for ordinary work — drafting, summarising, fixing a spreadsheet, with a person reading the result — Claude Sonnet 5.5 is about as safe as the model it replaces. Anthropic’s own testing says it crosses no new risk threshold. What decides your privacy is not the model at all. It is the plan you use it on and one setting in your account, and those did not change on launch day.

This post separates the two questions people run together. First, what the Claude Sonnet 5.5 system card says about the model’s behaviour. Second, what happens to what you type into it, which the system card does not cover and was never meant to.

What Claude Sonnet 5.5 is

Anthropic released Claude Sonnet 5.5 on 28 September 2026 as the second model in its 5.5 family, after Opus 5.5. It is the mid-priced model: Anthropic pitches it at well-scoped everyday work — fixing bugs, and producing documents, slides and spreadsheets — and describes it as more than 30% faster than Sonnet 5. The API price stayed at $2 per million input tokens and $10 per million output tokens, in US dollars.

It is in all the Claude apps and on the big clouds. AWS announced it on Amazon Bedrock the same day, and 9to5Mac covered the app update. If the difference between a model and the product you open is fuzzy, AI agent vs LLM sorts it out in a few minutes — and it matters here, because safety lives partly in each.

What the system card says about the model

Anthropic measures its models against a set of capability thresholds in its Responsible Scaling Policy — lines that, once crossed, require stronger safeguards. The system card reports that Sonnet 5.5 is broadly less capable than Opus 5.5 and does not cross any new threshold. It rates the risk of misalignment as low, partly because the model found it hard to control its own chain of thought or slip past monitors when its reasoning was examined.

That last point is worth comparing. When OpenAI published its card for GPT-6 Astra, it reported the opposite trend — reasoning that had become harder to watch. We went through that in is GPT-6 Astra safe. A model whose reasoning stays readable is easier to supervise, and here the vendor says that is still true.

On alignment more broadly, Anthropic’s announcement says Sonnet 5.5 improves on or matches Sonnet 5 on most measures. Most is not all. The card is where the exceptions live, and a vendor grading its own model is still a vendor grading its own model.

The two safeguards that are new

The first is about cybersecurity. Sonnet 5.5’s cyber skills are strong enough that Anthropic gave it the kind of limits it previously kept for its top models. According to The Next Web’s report on the launch, higher-risk security requests visibly fall back to Sonnet 5 instead of being answered by the new model. Security teams who need the full capability can apply to a verification programme for tiered access. For most people this is invisible. If you do security work, it explains why an answer suddenly came from the older model.

The second is newer and more interesting for ordinary users. Sonnet 5.5 is the first Sonnet to launch with classifiers that block attempts to extract its reasoning. The Next Web ties this to an incident in August 2026, when researchers decoded 315,320 thinking blocks from 6,708 public agent traces and recovered 62 API keys, 33 passwords and seven private keys. Read that list again. Those were secrets people had pasted into agent sessions, sitting in the model’s intermediate reasoning.

What happens to your data

None of this changed with Sonnet 5.5, which is why it is easy to miss. Data handling follows the account, not the model. On Claude Free, Pro and Max — including Claude Code used from those accounts — Anthropic’s consumer terms let you choose whether your chats are used to train future models. If you allow it, those chats are kept for up to five years. If you do not, the retention period is 30 days.

The same page says those consumer terms do not apply to Claude for Work, Claude for Government, Claude for Education, or API use, including through Amazon Bedrock and Google Cloud’s Vertex AI. Those run under commercial terms. Anthropic’s announcement also says Sonnet 5.5 is available with zero data retention, which is an arrangement on the business side, not a switch in the consumer app.

To check the consumer setting, open Settings, then Privacy, and look for the toggle Anthropic calls Help improve our AI models. Anthropic’s privacy centre spells out what switching it off does and does not do: new chats stop going into future training, but data already used in a training run stays in that model, and conversations flagged by safety classifiers can still be used for trust and safety work.

If you are in the European Economic Area, the UK or Switzerland, the data controller for your account is Anthropic Ireland, Limited, which is who a GDPR request goes to. The same question — where does this file go once I upload it — is the first thing uploading Excel to ChatGPT deals with, and the answer has the same shape for any assistant.

What this changes for using it at work

Very little, if your habits were already sound. A faster, cheaper model mostly means more people will hand it bigger jobs — whole folders, long agent runs, spreadsheets with real figures in them. That is where the practical risk sits, and it is the same risk as last week.

  1. Check which plan you are on before you paste anything confidential. A personal Pro account and a company Claude for Work account are governed by different terms, even when the model is identical.
  2. Look at the training toggle once, deliberately. Whichever way you set it, set it on purpose.
  3. Keep secrets out of chats and agent sessions. API keys, passwords and private keys belong in a password manager, not a prompt — the August incident is the reason.
  4. Check the work, not the confidence. A fluent wrong answer is the normal failure, and a faster model produces it faster.
  5. For coding, review the diff before you merge. Muse Code vs Claude Code walks through what that review should look for.

If you are choosing between assistants rather than models, Muse Spark vs Claude compares them plan by plan. For how a sibling vendor handled the same data questions this month, is GPT-6 Sol safe reads OpenAI’s version. And the cheapest way to get a better answer from any of them is still writing a prompt that works first try.

Coursium is a mobile app that teaches people to use AI at work, and it is on the App Store. Model names change every few weeks. Knowing where your data goes, keeping secrets out of prompts and checking output carry over from one to the next. If that is what you want to practise, have a look at Coursium.

Frequently asked questions

Is Claude Sonnet 5.5 safe to use for normal work?

For drafting, summarising, coding and document work with a person reviewing the result, yes. Anthropic’s system card reports that Sonnet 5.5 crosses no new threshold in its Responsible Scaling Policy and rates misalignment risk as low. As with any model, it can be confidently wrong, so the output still needs checking.

Does Anthropic train on my Claude Sonnet 5.5 chats?

On Claude Free, Pro and Max, that depends on a setting you choose. If you allow training, chats are kept for up to five years; if not, retention is 30 days. Claude for Work, Education, Government and API use, including Amazon Bedrock and Vertex AI, run under commercial terms that the consumer training setting does not cover.

Why do some security questions get answered by Sonnet 5 instead?

Anthropic applies cyber safeguards to Sonnet 5.5 similar to those on its top models. Higher-risk cybersecurity requests visibly fall back to Sonnet 5. Security teams that need the full capability can apply to Anthropic’s verification programme for tiered access.

Who is the data controller for Claude in Europe?

For users in the European Economic Area, the UK and Switzerland, Anthropic’s privacy centre names Anthropic Ireland, Limited as the data controller. That is the entity a GDPR access or deletion request is addressed to.

Coursium

Stay ahead of AI — learn the tools on your phone.

Get the app