Skip to content
← BACK TO BLOG
Fikri Firman Fadilah
6 min read
Engineering

Why Cheaper Models Can Win on Cost per Completed Task

Share:
Share on Twitter
Share on LinkedIn
Copy Link

How to compare LLMs by cost per completed task, and route work between Haiku and Sol with validated escalation.

Last quarter, a planning doc proposed moving our entire extraction pipeline to the newest flagship model because its benchmark numbers were higher. The pipeline makes tens of thousands of short calls a night. Before anyone approved it, I wanted a different number: what does one finished task cost us?

That question is the core of this post. The short version is that a more expensive model can be the wrong choice for most of your traffic and the right one for a narrow slice of it.

Start with the workload, not the leaderboardLink to this section

The premise that Haiku 5.5 beats GPT-6.1 Sol is too broad to act on. In the cited Terminal-Bench runs, Sol 6.1 leads Haiku by 19 points on difficult terminal coding [1]. Haiku also trails Sonnet 5.5 on that same demanding evaluation [2]. If your workload is hard command-line work, Sol is the better-supported choice among these models.

Haiku's case is different. Its strongest argument is volume: extracting facts, classifying requests, preparing summaries, and handling short agent tasks that repeat thousands of times [2]. Those are the jobs where a small model's per-call economics matter and where the work rarely needs deep multi-step reasoning. The decision starts by splitting your traffic into those two buckets.

Cost per completed task, not cost per tokenLink to this section

Sol 6.1 is priced at $2 per million input tokens and $10 per million output tokens [1]. Haiku's short-prompt rates match GPT-6 Luna, and by the source's own arithmetic Sol's rate is about twenty times the small models' short-prompt rates [2][1]. On a rate card alone, the answer is obvious.

The rate card is the wrong unit. What you pay is what it costs to finish a task, and that depends on how many tokens a model spends getting there. Artificial Analysis puts GPT-6.1 Sol at $0.72 per task on its Intelligence Index [3]. Its figure for GPT-6 Luna is $0.07 per task, against $0.21 for Haiku at max effort [4]. Luna and Haiku share sticker rates, yet Haiku costs about three times as much per task. Verbosity and effort settings drive that gap. Haiku's default effort is medium, not max, so the configuration you launch with changes the bill [4].

The metric I use is simple:

cost_per_completed = total_spend_including_retries_and_escalations / tasks_that_passed_validation

Note the denominator. A cheap model that fails 30 percent of the time is not cheap if you pay for the failures and then pay again to escalate them. Here are illustrative numbers, not measurements:

StrategySpend per 1,000 tasksTasks completedCost per completed task
Sol only$5.00970$0.0052
Haiku first, Sol on failure$1.401,000$0.0014

The routed version wins only if the first-pass completion rate is high and your validator catches failures reliably. Change either assumption and the result changes.

Integration is a cost tooLink to this section

Before writing any routing code, check the integration path. In our notes, Sol 6.1 requires the Responses API for tool calling, unlike the Chat Completions models we already use. Verify that against the current provider documentation before building on it, because it changes the shape of your code.

In practice, this means two client paths, two tool-call output formats, and two places where retries and logging need to match. Put both behind one adapter that returns a normalized result. Otherwise your router ends up full of provider conditionals, and the escalation logic becomes hard to test.

Evidence limitsLink to this section

Treat the cross-model numbers as directional. The cited runs used different agents and execution policies, so the 19-point gap and Haiku's 22.8-point lead over Luna are not pure measurements of the models under identical conditions [1]. Every figure in the comparison is vendor-reported, with competitor results taken from public reports [5]. A "not reported" cell means the source has no number, not zero [2].

Community praise for Haiku's coding is also anecdotal. It may be accurate for your stack, but it does not tell you anything about your failure rate. The only evidence that matters for your decision is a small eval on your own tasks, with your prompts, your effort setting, and your validator.

Recommendation: route by task type and escalate on failureLink to this section

Send extraction and classification to Haiku. Send hard terminal debugging and multi-file work to Sol directly. Escalate from Haiku to Sol only when the output fails a check. Here is a starting config:

yaml
routing:
  routes:
    extract_fields: fast
    classify_request: fast
    summarize_short: fast
    terminal_debug: strong
    multi_file_edit: strong
  tiers:
    fast:
      model: claude-haiku-5.5   # confirm the exact model ID
      api: chat
      effort: medium
      escalates_to: strong
    strong:
      model: gpt-6.1-sol        # confirm the exact model ID
      api: responses
      escalates_to: null
  max_escalations: 1

The router itself is short:

python
import jsonschema

def check(task, output):
    if output is None:
        return "malformed_tool_call"
    if not output.get("complete", False):
        return "unfinished"
    try:
        jsonschema.validate(output, SCHEMAS[task.kind])
    except jsonschema.ValidationError:
        return "schema_invalid"
    return None

def escalation_path(start, tiers, max_escalations):
    path = [tiers[start]]
    while len(path) <= max_escalations and path[-1].escalates_to:
        path.append(tiers[path[-1].escalates_to])
    return path

def run_task(task, route_table, tiers, call, log):
    spent = 0.0
    for tier in escalation_path(route_table[task.kind], tiers, max_escalations=1):
        output, cost = call(tier, task)  # output is None if tool args don't parse
        spent += cost
        reason = check(task, output)
        if reason is None:
            return {"status": "completed", "tier": tier.name, "spent": spent, "output": output}
        log.info("escalating", task_id=task.id, tier=tier.name, reason=reason)
    return {"status": "failed", "spent": spent, "output": None}

Log the tier, reason, and cost for every attempt. Without that, you cannot compute cost per completed task after the fact.

Failure modes to watchLink to this section

  • Unfinished tasks. A response can be schema-valid and still incomplete, such as an extraction that returns three of five required fields. Schema validation alone won't catch this. Require an explicit completion flag or a minimum count of required fields.
  • Malformed tool calls. Arguments that don't parse or name a tool that doesn't exist are not worth retrying on the same tier. Escalate them instead, and count them separately so you can see whether a prompt change caused them.
  • A validator that is too strict. If the check rejects good outputs, you pay for Haiku and Sol on the same task. Watch cost per completed task, not just the escalation rate.
  • Escalation drift. If the escalation rate climbs from around 12 percent to 40 percent, something changed. It might be a prompt edit, a new input distribution, or a silent effort change. Alert on it.
  • Effort creep. Someone sets max effort "just for accuracy." That can erase the savings without any visible change in the code.

When not to do thisLink to this section

Skip routing if your volume is low. The complexity of two providers and a validator costs more engineering time than it saves in model spend. Also avoid cheap-first routing for long-horizon coding, where a failed first pass wastes many steps, and for any task where a wrong answer is costly and you cannot validate it cheaply.

TakeawaysLink to this section

  • Split your traffic by task type before comparing models. Haiku and Sol are candidates for different buckets, not competitors for all of them.
  • Measure cost per completed task, including retries and escalations. Sticker prices and per-task averages from leaderboards won't tell you this.
  • Confirm the Sol integration path early. Tool calling through the Responses API changes your client code and should be behind a single adapter.
  • Build a small eval on your own tasks, and treat published cross-model gaps as directional, since the runs used different agents and policies.
  • Make the validator strict on completion, not just on schema, and alert when the escalation rate moves.

SourcesLink to this section

  1. GPT Sol & Luna vs Claude Haiku 5.5: Specs and Benchmarks - Kingy AI
  2. Claude Haiku 5.5: Benchmarks, Specs, Sol & Luna Compared
  3. LLM Cost per Task: Same $2/$10 Price, 10x Different Bills
  4. Claude Haiku 5.5 vs GPT-6 Luna: Same Price, Same Bill?
  5. GPT-6.1 Sol vs Claude Sonnet 5.5: Benchmarks, Pricing, and a Hands-On Test

Recommendations

You might also like