Claude Haiku 5.5
Claude Haiku 5.5 is our fastest, most capable small model.
Claude Haiku 5.5 is our fastest, most capable small model. Built for high-volume work like summarization, subagents, and browser use. Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model weโve ever released. Claude Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles quick and repetitive workloads (like summaries, compactions, database queries, and classification requests). It pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work. And, since itโs also our fastest model to date, it works especially well for speed-sensitive tasks like live customer support and browser use.ยน Haiku 5.5 is available at a much lower price than Haiku 4.5. On average, it now costs around 75% less to run.ยฒ Hereโs how Claude Haiku 5.5 performs across a range of benchmarks: For details on how we run our evaluations, see the Haiku 5.5 System Card . Haiku 5.5 is our first Haiku-class model to come with an adjustable effort setting. This means that, as with our other models, users can decide whether to optimize for cost or intelligence. The charts below show how Haiku 5.5 performs on three benchmarks at each effort setting: OSWorld 2.1 measures how well agents can operate a real computer to finish long, multi-step tasks. Artificial Analysisโs GDPval-AA v2.1 evaluates agents on real-world professional work across 44 occupations. Humanityโs Last Exam (HLE) is a test of expert-level academic knowledge and reasoning. In early testing, our customers reported results consistent with the performance and cost improvements shown above.
Hereโs what they told us about the new model: โWeโre very impressed with Claude Haiku 5.5, particularly its speed. We ran it through our eval suite for AI Teammates, our AI agent product, covering use cases like triaging bugs, setting up projects, and searching large portfolios to surface high-risk or overdue work. Compared with the model we use today, we saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn. Itโs a noticeably snappier experience.โ โAt Hub Spot, we use simulated portals to evaluate new models on CRM tasks like reporting on deals. We mostly test the smaller, more efficient models, and Claude Haiku 5.5 got the best score weโve seen on this suite yet, at 92.8% averaged over three runs. One CRM audit task asks models to identify stale but ambiguous records. Across all of the models we tested, Haiku 5.5 was fastest to complete the task, and had the highest hit rate and the lowest false positive rate.โ โAsk in Document is one of our big sources of spend, doing about 8M calls a week in production. It answers very specific questions on top of one or a few documents. We ran 400 queries, and Claude Haiku 5.5 was a statistically significant improvement over Haiku 4.5: 0.84 vs. 0.76.โ โOur customers use Box AI across large volumes of their enterprise content. With widespread usage comes the need to manage efficiency and cost, and to find the best model to suit the task at hand. In early testing, Claude Haiku 5.5 scored 11 points higher than Haiku 4.5 at about half the latency. Weโd put it to use on analytical work that runs at scale, from cost reports to financial summaries and weekly recurring reviews.โ โThe short and high-volume work is where Claude Haiku 5.5 fits for us, like quick lookups, subagents, and summaries. While a bigger model builds the deck, a Haiku 5.5 subagent goes into the 10-K and pulls the segment revenue line the deck needs.
Itโs accurate enough that weโd trust it there, and fast and cheap enough that we can run it a lot.โ โClaude Haiku 5.5 joins the sidekick lineup in Devin Fusion as an excellent option. With Haiku 5.5 as the sidekick, Fusion holds a top-tier Frontier Code score of 66.2 while cutting cost and latency. You can try it today in the Devin CLI with Opus 5.5 as the lead.โ The table below shows how Claude Haiku 5.5โs pricing compares to our other models. Haiku 5.5 is especially good value when used for tasks with prompts up to 100,000 tokens, which make up around 90% of requests to our previous Haiku model. Alignment. Claude Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to Haiku 4.5. In particular, we found far fewer instances of misaligned behavior, and a lower willingness to cooperate with misuse. The modelโs system card describes our evaluation process and results in more detail. Safeguards. Consistent with its capabilities, Haiku 5.5โs cybersecurity safeguards are more restrictive than Haiku 4.5โs, but somewhat less restrictive than those weโve applied to other recent models. In cybersecurity, they permit a wider range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers. Haiku 5.5โs biology safeguards are the same as for Sonnet 5, Sonnet 5.5, and Opus 5. They allow research biology questions but restrict access to requests that we judge as likely to cause harm.
โก Effects Interpreter
๐World Economy
- โถExport-heavy economies may see demand wobble as buyers wait and watch.
- โถThe ripples can spread across borders, nudging growth forecasts here and there.
๐๏ธLocal Economy
- โถLocal wages and hours worked might bend slightly with the wider trend.
- โถEveryday costs in your town may drift as the wider economy reacts.
๐ฆRates & Banks
- โถLenders usually hold their nerve until a clearer trend appears.
- โถRate decisions tend to follow data, not headlines, so patience is the norm.
โค๏ธHealth
- โถLocal wellbeing services are there precisely for moments when news feels overwhelming.
- โถTalking things through with family or friends can ease the load.
๐ทWealth
- โถYour pension or investments might sway a touch as markets digest this.
- โถKeeping perspective on your timeline usually beats reacting to any single story.
๐ Housing
- โถBuyers and renters may notice only a gentle drift, if anything at all.
- โถEstate agents usually say the market takes weeks to catch up with news.