Skip to content
← BACK TO BLOG
Fikri Firman Fadilah
1 min read
Engineering

Running Claude Haiku 5.5 for High-Volume Ticket Triage

Share:
Share on Twitter
Share on LinkedIn
Copy Link

What Claude Haiku 5.5 costs, what its reported benchmarks show, and how to call it from Python for cheap, high-volume classification.

Our support queue hit about 50,000 tickets a day last quarter, and most of them needed nothing more than a label and an urgency flag. We were running them through a large model because it was the one we trusted, and the bill showed it. When Anthropic released Claude Haiku 5.5 on October 7, 2026, it looked like the obvious candidate for that kind of work [3][4].

This post covers what the model is, what the published numbers do and do not tell you, and a small Python script for triage. I also cover where I would not use it.

What Claude Haiku 5.5 is

Haiku 5.5 is the third model in the Claude 5.5 family, after Opus 5.5 on September 22 and Sonnet 5.5 on September 28 [1]. Anthropic positions it for high-volume, cost-sensitive work: classification, summarization, extraction, live support, voice agents, and in-app assistants [3]. Anthropic also describes it as the successor to Haiku 4.5 and says it costs 75% less than that predecessor [3]. One video summary reports a 90% reduction [5], so treat the exact figure as something to verify against your own traffic.

Three details matter more for engineering decisions than the headline:

  • Context window: 1M tokens, with text and image input [2][4].
  • Adjustable effort: OpenRouter describes it as the first Haiku with adjustable effort. Thinking is adaptive and on by default, and you can turn it off at low, medium, and high effort [4].
  • Built-in safeguards: Anthropic says this is its first Haiku with safeguards for a narrow set of high-risk cybersecurity requests, and that most everyday tasks should be unaffected [3].

The safeguards are the one I would test first. If your product touches security tooling, pentest reporting, or anything that looks like exploit discussion, you may get refusals that the larger models did not produce for you.

Pricing and the cost math

Pricing for prompts under 100,000 tokens is $0.10 per million input tokens and $0.50 per million output tokens. Prompts over 100K tokens rise to $0.50 and $2.50 [3][4]. LLM Stats lists cached input at $0.010 per million [2], which matters if your system prompt is long and identical across requests.

Here is the back-of-envelope estimate for our triage workload: 50,000 requests a day, each with about 2,000 input tokens and 300 output tokens.

  • Input: 100M tokens × $0.10/M = $10.00/day
  • Output: 15M tokens × $0.50/M = $7.50/day
  • Total: about $17.50/day before caching

Two cautions. First, the output price applies to any reasoning tokens the model generates, not just the visible JSON. Thinking is on by default [4], so a classifier that quietly reasons for 800 tokens costs far more than one that answers in 20. Second, these numbers exclude retries, and I have not measured them.

What the benchmarks do and do not tell you

Anthropic's reported numbers include OSWorld 2.1, a computer-use benchmark, where it reports 72.4% for Haiku 5.5 against 15.7% for Haiku 4.5 [5]. That is a large jump, and it is the number most likely to shape how people think about the model.

I would not read it as a triage benchmark. Computer use is a different task from picking one of five labels for a support ticket. The video source also notes that all benchmark numbers are Anthropic's own and that independent results were not yet available at the time [5]. LLM Stats has a leaderboard view, but its ranking depends on how its composite score is weighted, so I would not use it as a single number either [2].

The benchmark I trust is the one I run on my own labeled tickets. Anything else is a prompt for your own evaluation, not a substitute for one.

A simple triage tutorial in Python

OpenRouter exposes an OpenAI-compatible endpoint, so you can use the standard openai SDK and change only the base URL and model slug [4]. Install the SDK and set your key:

bash
pip install openai
export OPENROUTER_API_KEY="sk-or-..."

Next, the classifier. It asks for JSON with one of five labels and an urgency flag:

python
import json
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
)

LABELS = ["billing", "bug", "feature-request", "account", "other"]

SYSTEM = (
    "Classify the support ticket. Reply with JSON only, in this shape: "
    '{"label": one of ' + json.dumps(LABELS) + ', "urgent": true or false}.'
)

def triage(ticket_text: str) -> dict:
    resp = client.chat.completions.create(
        model="anthropic/claude-haiku-5.5",
        messages=[
            {"role": "system", "content": SYSTEM},
            {"role": "user", "content": ticket_text},
        ],
        temperature=0,
        max_tokens=300,  # leave room for any reasoning tokens
    )
    raw = resp.choices[0].message.content or ""
    try:
        result = json.loads(raw)
    except json.JSONDecodeError:
        return {"label": "other", "urgent": True, "error": raw[:200]}

    if result.get("label") not in LABELS:
        result["label"] = "other"
    result["_usage"] = resp.usage.total_tokens
    return result

if __name__ == "__main__":
    print(triage("I was charged twice for March and I need a refund today."))

Three choices here are deliberate. temperature=0 makes runs more repeatable, though it does not guarantee identical output. The fallback sends unparseable output to a human-review label instead of crashing the worker. And max_tokens is set well above the JSON size, so reasoning tokens do not truncate the answer.

Before you ship, run it against a labeled sample. A minimal eval loop looks like this:

python
import csv
from collections import Counter

def evaluate(path: str) -> None:
    correct, total, tokens = 0, 0, 0
Share:
Share on Twitter
Share on LinkedIn
Copy Link

Recommendations

You might also like