Skip to navigation

span-router

One model ID that picks the model for each task, with context, capabilities, and cache-aware cost in mind.

span-router puts model selection inside the gateway. Powered by span-01, it combines task context, model capabilities, and cache-aware cost estimates behind one model ID: span-router.

You pay the serving model’s list price, including applicable cache discounts. No markup.

To send your first request, see the Quickstart.

Route the task, not just the prompt

“Fix this” tells the router very little. “Fix this failing migration in a production database; keep existing records and use the database tools” tells it much more.

Your application already knows that context. Put it in the system or developer message: what the agent does, where it runs, and the constraints it must respect.

span-router reads that context alongside the latest exchange to classify the task. It also checks that a model can handle the request’s context length, images, tools, and structured-output requirements.

Keep continuity across turns

Send the same thread_identifier on every turn. span-router prefers to keep the current model while it remains usable, and rechecks capabilities and context limits as the conversation grows. When the cost-aware switching policy is on, it can also reconsider the route using cost estimates.

A new turn alone isn’t a reason to switch.

Cache-aware cost

The cheaper model isn’t always the cheaper next call. A model with a warm cache can cost less than moving a long conversation to a lower-priced model that reads it from scratch.

The cost-aware policy weighs cached reads, applicable cache writes, fresh input, and expected output. List price is one input, not the whole decision.

To get the most from caching:

  • Keep system instructions and tool definitions stable.
  • Send the full conversation history on every turn.

Cache reuse depends on the provider and on matching prompt prefixes, so it isn’t guaranteed.

Benchmarks

Single requests

On 557 held-out requests, span-router passed 94.1% at 37% lower cost than the most accurate fixed model.

ConfigurationPass rateCost
span-router94.1%$0.46
Most accurate single model94.3%$0.73
Cheapest single model91.0%$0.09
Best model per item, in hindsight (upper bound)97.1%$0.18

The items come from HumanEval, MBPP, GSM8K, MATH-500, MMLU-Pro, BFCL, LiveCodeBench v6, and an internal set. None were used to tune routing, and every item was run on every candidate model. The hindsight row picks a model after seeing each outcome, so it isn’t a deployable router. A separate live run through span-router scored 93.0%.

Agent tasks

Benchmarkspan-routerJev RouterOpenRouter Auto
τ-bench airline (50 tasks)0.8870.8330.460
τ-bench retail (114 tasks)0.9330.8860.623
τ³-bench banking (97 tasks)0.3130.3060.041

span-router passed 82.0% of airline tasks in all three trials, versus 72.0% for Jev Router. On retail, that was 86.8% versus 78.9%. Average airline task time was 28.6 seconds versus 32.9 seconds. The strongest gains are in airline and retail; banking is close.

span-router and Jev Router scores are averaged over three trials; OpenRouter Auto was run once. Every router used the same τ user simulator, GPT-4.1 at temperature 0. Measured September 26–27, 2026.

Call it

Use the Respan Chat Completions endpoint. Change the model name, put your agent’s context in the system message, and keep a stable thread identifier:

response = client.chat.completions.create(
model="span-router",
messages=[
{"role": "system", "content": "You are a support agent for an airline. Use the booking tools; never change a fare without confirmation."},
*conversation_history,
],
extra_body={"thread_identifier": "conversation_123"},
)
  • Streaming, tool calls, and structured outputs are supported.
  • Send the full message history with the same thread_identifier on every turn.
  • Don’t combine span-router with fallback or load-balancing options.
  • Calls use Respan credits. Your own provider keys aren’t supported.

Pricing

Each request is priced using the model that served it, with no router markup:

  • Uncached input and output use that model’s rates.
  • Cached input uses the model’s cache-read rate.
  • When a provider charges for cache writes, those tokens use its cache-write rate.

Check each request’s cost in its Respan log.