Claude Haiku 5.5 gets tighter cyber safeguards, lower price
Anthropic has released Claude Haiku 5.5, its cheaper and faster model for repetitive, speed-sensitive work. The company says it finds vulnerabilities and writes exploits better than Haiku 4.5, so it has added stricter cybersecurity safeguards. Its offensive skills still trail well behind Anthropic's larger models.
The company also describes Haiku 5.5 as its most resistant Haiku model yet to prompt injection, where malicious instructions are hidden in content the AI reads, such as an email or a webpage.
Offensive skills measured without guardrails
Anthropic tested the model's offensive capabilities with its cybersecurity safeguards turned off. In one test built around known flaws in Chrome's V8 JavaScript engine, Haiku 5.5 reached arbitrary code execution in four out of 410 runs.
A separate evaluation looked at multi-stage cyber operations. Here a pre-release build completed 3.3% of challenges. Sonnet 5.5 managed 46.1% and Opus 5.5 reached 67.6%. On the ExploitGym benchmark, Haiku 5.5 beat Claude Sonnet 5 but stayed far below the rest of the 5.5 family.
Anthropic notes that these numbers do not necessarily show what ordinary users can do with the models once general-access safeguards are in place.
The new safeguards are stricter than those on Haiku 4.5 but somewhat looser than on other recent Anthropic models. They allow more defensive work than Sonnet 5.5's restrictions, while still blocking penetration testing and other techniques that attackers are more likely to use. Security professionals who qualify can request fewer restrictions through the company's Cyber Verification Program.
Fewer wrong refusals
Anthropic ran the model against harmful requests, harmless questions on sensitive topics, and conversations in which a simulated user slowly pushed it toward a harmful result. The topics included weapons, extremism, tracking and surveillance, child safety, mental health and election integrity.
With a near-final version of the Claude.ai production system prompt applied, Haiku 5.5 gave harmless responses to 99.71% of harmful requests. It wrongly refused 0.82% of harmless requests, compared with 3.05% for Haiku 4.5.
In longer conversations it did better on influence operations, tracking and surveillance. It did worse on weapons scenarios when tested through the API without a system prompt. Anthropic also flagged remaining weak spots in sensitive conversations, including self-harm and eating disorders, and tells API developers to add their own safeguards. These tests did not include extra production protections such as real-time probes and monitoring.
Coding and computer-use misuse
In Claude Code tests, Haiku 5.5 refused 84.3% of malicious requests, up from 66.6% for its predecessor. The requests included writing malware, supporting DDoS attacks and building non-consensual monitoring software. At the same time, it helped with more permitted security tasks, such as analyzing penetration-test results.
In computer-use tests covering surveillance, unauthorized data collection and similar activity, it refused about 82.6% of requests, up from 58.9%. That was also a higher refusal rate than Sonnet 5.5 and Opus 5.5 achieved on the same evaluation.
Against adaptive attackers in coding and computer-use environments, its prompt-injection resistance largely matched Anthropic's frontier models. On a separate Gray Swan benchmark, though, it was less resistant than Sonnet 5.5 and Opus 5.5, with most of the remaining weakness in graphical computer use. Most tests excluded extra production safeguards, while the adaptive-attack results were reported with and without prompt-injection probes.
Pricing and availability
Haiku 5.5 is cheaper than Haiku 4.5. For prompts of up to 100,000 tokens, which Anthropic says covered most Haiku 4.5 requests, input and output token prices are 90% lower. An adjustable effort setting lets users trade cost against intelligence, and the model can act as a coding subagent for Opus 5.5 and Sonnet 5.5.
Aaron Vinh, Staff Software Engineer at Asana, said that in tests for the company's AI Teammates agent product, they "saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn."
The model is available on all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure, under the identifier claude-haiku-5-5. Anthropic is also halving cache read prices for Sonnet 5.5, which it says makes most agent tasks around 20% cheaper. Max and Team subscribers will get monthly API credits, and the Python and TypeScript SDKs gain beta support for computer and browser use.
Our Take
A model that is both cheaper and better at exploitation will probably be used in more automated pipelines, so the tighter cyber restrictions look like a sensible trade-off. The jump in refusal rates for coding and computer-use misuse matters because these are the settings where agents act on their own. Prompt injection remains the harder problem. Graphical computer use is still the weak spot, and the beta browser and computer-use support in the SDKs could widen exposure. Developers building agents should take Anthropic's advice and add their own controls, including adversarial testing of AI agents. It is worth watching whether independent testing confirms Anthropic's results once production safeguards are switched on.
