top of page

Anthropic launches Claude Haiku 5.5 with 75% lower task costs, faster inference and adjustable reasoning

2 minutes ago
6 min read
Anthropic launches Claude Haiku 5.5 with 75% lower task costs, faster inference and adjustable reasoning

Anthropic has introduced Claude Haiku 5.5, a smaller model in the Claude 5.5 family designed for applications where inference speed, operating cost and the ability to process large numbers of requests are critical.


Announced on October 7, 2026, the model combines faster inference with adjustable reasoning, allowing developers to control how much computational effort the model dedicates to individual requests.


Anthropic reports up to 75% lower costs per completed task in its evaluations. This is a task-level efficiency claim rather than a universal reduction in API token prices: the actual saving depends on the workload, reasoning settings, token consumption and success rate.


Claude Haiku 5.5 is positioned for high-volume applications, including customer support, information extraction, classification, coding assistance and AI agents that execute many relatively short operations.


··········


CLAUDE HAIKU 5.5 AT A GLANCE


........


Specification

Claude Haiku 5.5

Developer

Anthropic

Announcement date

October 7, 2026

Model family

Claude 5.5

Model positioning

Fast, cost-efficient inference

Reported task-cost reduction

Up to 75%

Reasoning

Adjustable

API input pricing

From $0.10 per million tokens

API output pricing

From $0.50 per million tokens

Pricing qualification

Published starting rates for prompts up to 100K tokens

Main workloads

Classification, extraction, support, coding and agents

Primary optimization

Cost and latency per completed task


........


The published starting prices and Anthropic's task-cost claims describe different aspects of the model's economics. API rates establish the charge per token, while task-level cost also depends on how many tokens and attempts are required to produce an acceptable result.


··········


WHY A 75% REDUCTION IN TASK COST IS DIFFERENT FROM CHEAPER TOKENS


Anthropic's headline efficiency claim concerns the cost of completing a task, which is a more useful metric for many production applications than the price of an individual token.


A model with inexpensive output tokens can still become costly if it generates unnecessarily long responses, repeatedly invokes tools or fails often enough to require retries.


Conversely, a model with higher nominal token prices may complete some workflows more economically if it requires fewer steps.


For an agentic application, the relevant calculation is:


Cost per successful task = total inference and tool-related costs ÷ successfully completed tasks


An illustrative Data Studios calculation shows the economic effect of Anthropic's reported reduction.


If a workload currently costs $0.20 per completed task, a 75% reduction would bring that cost to $0.05, assuming comparable task definitions, quality requirements and measurement conditions.


........


Monthly completed tasks

At $0.20/task

At $0.05/task

Illustrative saving

10,000

$2,000

$500

$1,500

100,000

$20,000

$5,000

$15,000

1,000,000

$200,000

$50,000

$150,000


........


These figures are Data Studios scenarios, not observed Haiku 5.5 customer results. They illustrate the effect of a 75% reduction if that improvement is achieved consistently at the application level.


The percentage should not be applied automatically to every workload. Developers need to measure total execution cost against an appropriate baseline while preserving equivalent output quality.


··········


ADJUSTABLE REASONING LETS DEVELOPERS CONTROL COMPUTATIONAL EFFORT


Claude Haiku 5.5 supports adjustable reasoning, allowing applications to allocate different levels of reasoning effort depending on the complexity of the request.


This addresses a common inefficiency in AI deployments: simple tasks and difficult tasks do not require the same amount of computation.


For example, extracting an invoice number from a well-structured document is different from resolving conflicting information across multiple records. Applying the same reasoning effort to both can increase latency and cost without producing a proportional quality improvement.


With adjustable reasoning, developers can configure a lighter execution path for routine operations and greater effort for requests requiring more complex analysis.


The relationship is not necessarily linear. More reasoning does not guarantee a better answer, and the optimal setting depends on the model, prompt and evaluation criteria.


Production systems should therefore test reasoning configurations against actual workload distributions rather than assuming that maximum effort produces the best economic result.


··········


FASTER INFERENCE CAN IMPROVE AGENT THROUGHPUT


Inference speed affects applications differently depending on whether model calls occur independently or sequentially.


In a conventional classification pipeline, multiple requests can often be processed concurrently. Higher throughput allows the application to complete more operations within a given time.


In an AI agent, model calls frequently depend on previous results. The agent may interpret a request, select a tool, examine its output and decide what to do next.


When these operations are sequential, latency accumulates.


A workflow involving ten dependent model calls, for example, cannot eliminate the latency of earlier calls simply by processing later steps in parallel.


A faster model can therefore reduce the duration of the entire workflow, provided that model inference represents a meaningful share of total execution time.


External API delays, tool execution, database queries and network communication can still dominate overall latency.


For developers, the useful performance indicators include time to first token, output generation speed, total request latency and end-to-end task completion time. These should be measured separately.


··········


BENCHMARKS SHOW LARGE GAINS OVER HAIKU 4.5


Anthropic's published evaluations show improvements across knowledge work, computer use, reasoning and agentic coding.


........


Benchmark

Haiku 4.5

Haiku 5.5

Sonnet 5.5

GDPval-AA v2.1

735

1,620

1,840

OSWorld 2.1, offline subset

15.7%

72.4%

83.9%

Humanity's Last Exam, no tools

10.2%

45.9%

56.9%

Terminal-Bench 4.0

0.0%

39.2%

70.6%

Chartography, no tools

6.4%

46.4%

61.6%


........


These are Anthropic-reported results, not independent Data Studios testing. The benchmarks measure different capabilities and should not be combined into a single overall performance score.


The 72.4% OSWorld result is particularly relevant for desktop automation, while Terminal-Bench shows that Haiku 5.5 has improved substantially but remains behind Sonnet 5.5 on complex agentic coding.


··········


EARLY ENTERPRISE TESTS PROVIDE REAL-WORKLOAD EVIDENCE


Anthropic also published results from early enterprise evaluations.


Asana reported over 30% lower task-completion latency and up to 2.5 times faster inference per agent turn compared with the model it currently uses.


HubSpot reported a 92.8% average score across three runs of its simulated CRM evaluation, its highest result among the models tested.


Box reported an 11-point improvement over Haiku 4.5 at approximately half the latency, while AlphaSense observed an improvement from 0.76 to 0.84 on a 400-query document-answering evaluation.


These results are useful because they involve specific production-like workloads, although the tests differ in baselines, datasets and evaluation methods. They do not establish a universal percentage improvement for every customer.


··········


API PRICING DEPENDS ON PROMPT LENGTH


Haiku 5.5 introduces two pricing bands based on prompt length.


........


API charge per 1M tokens

Up to 100K prompt

Over 100K prompt

Input

$0.10

$0.50

Output

$0.50

$2.50

Cache reads

$0.01

$0.05

Cache writes

$0.125

$0.625


........


For prompts up to 100,000 tokens, these input and output prices are 90% below Haiku 4.5's corresponding token rates. Above that threshold, the reduction is 50%.


Anthropic says approximately 90% of requests to Haiku 4.5 fell within the shorter-prompt category. It also notes that Haiku 5.5's updated tokenizer can consume slightly more tokens for an equivalent task, helping explain why its average task-cost reduction is approximately 75%, rather than 90%.


This makes prompt length, caching behavior and actual token consumption important variables when estimating migration savings.


··········


SMALLER MODELS CAN HANDLE SUBTASKS WHILE LARGER MODELS CONTROL THE WORKFLOW


Haiku 5.5 is particularly suited to architectures where a larger reasoning model coordinates multiple smaller agents.


A coding agent, for example, can delegate repository searches, log summarization, document extraction or narrowly defined code changes to Haiku while retaining Sonnet or Opus for difficult planning and verification.


The economic advantage depends on assigning work appropriately. Routing a complex task to a smaller model can erase token savings if it produces errors, requires repeated attempts or needs escalation.


The model is also available for browser and desktop automation, where rapid decisions and repeated tool interactions can make latency and cost especially important.


··········


AVAILABILITY AND ADDITIONAL CLAUDE PRICING CHANGES


Claude Haiku 5.5 is available through Claude, Claude Code, the Claude Platform and supported cloud providers, including AWS, Google Cloud and Microsoft Azure.


Developers can access it using the model identifier claude-haiku-5-5.


Alongside the launch, Anthropic has reduced Claude Sonnet 5.5 cache-read pricing by 50%, from $0.20 to $0.10 per million tokens. The company estimates this lowers costs for many Sonnet-based agentic workloads by approximately 20%.


Anthropic has also announced monthly Claude Platform API credits for eligible Max and Team subscribers, supporting experimentation with applications and agents.


··········


HAIKU 5.5 EXPANDS THE ECONOMICS OF HIGH-VOLUME AI


Claude Haiku 5.5 combines a substantial reduction in token pricing with stronger benchmark performance, adjustable reasoning and faster inference.


Its most promising deployment scenarios are workloads involving frequent, well-defined operations where speed and cost accumulate across thousands or millions of requests.


For complex coding and demanding reasoning, Anthropic's own evaluations still favor larger Claude models. For retrieval-assisted tasks, customer support, document processing, browser automation and subordinate agents, Haiku 5.5 offers a lower-cost execution layer.


The relevant production decision is therefore which tasks can move to Haiku 5.5 without compromising completion quality, and how much end-to-end cost and latency those migrations actually eliminate.


··········


FOLLOW US FOR MORE.


DATA STUDIOS


datastudios.org

bottom of page