> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://respan.ai/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://respan.ai/docs/_mcp/server.

# span-router

span-router puts model selection inside the gateway. Powered by [span-01](/docs/documentation/span-01/concept), it combines task context, model capabilities, and cache-aware cost estimates behind one model ID: `span-router`.

You pay the serving model's list price, including applicable cache discounts. No markup.

To send your first request, see the [Quickstart](/docs/documentation/features/gateway/span-router/quickstart).

## Route the task, not just the prompt

"Fix this" tells the router very little. "Fix this failing migration in a production database; keep existing records and use the database tools" tells it much more.

Your application already knows that context. Put it in the system or developer message: what the agent does, where it runs, and the constraints it must respect.

span-router reads that context alongside the latest exchange to classify the task. It also checks that a model can handle the request's context length, images, tools, and structured-output requirements.

## Keep continuity across turns

Send the same `thread_identifier` on every turn. span-router prefers to keep the current model while it remains usable, and rechecks capabilities and context limits as the conversation grows. When the cost-aware switching policy is on, it can also reconsider the route using cost estimates.

A new turn alone isn't a reason to switch.

## Cache-aware cost

The cheaper model isn't always the cheaper next call. A model with a warm cache can cost less than moving a long conversation to a lower-priced model that reads it from scratch.

The cost-aware policy weighs cached reads, applicable cache writes, fresh input, and expected output. List price is one input, not the whole decision.

To get the most from caching:

* Keep system instructions and tool definitions stable.
* Send the full conversation history on every turn.

Cache reuse depends on the provider and on matching prompt prefixes, so it isn't guaranteed.

## Benchmarks

### Single requests

On 557 held-out requests, span-router passed 94.1% at 37% lower cost than the most accurate fixed model.

| Configuration                                   | Pass rate | Cost       |
| ----------------------------------------------- | --------- | ---------- |
| **span-router**                                 | **94.1%** | **\$0.46** |
| Most accurate single model                      | 94.3%     | \$0.73     |
| Cheapest single model                           | 91.0%     | \$0.09     |
| Best model per item, in hindsight (upper bound) | 97.1%     | \$0.18     |

The items come from HumanEval, MBPP, GSM8K, MATH-500, MMLU-Pro, BFCL, LiveCodeBench v6, and an internal set. None were used to tune routing, and every item was run on every candidate model. The hindsight row picks a model after seeing each outcome, so it isn't a deployable router. A separate live run through span-router scored 93.0%.

### Agent tasks

| Benchmark                   | span-router | Jev Router | OpenRouter Auto |
| --------------------------- | ----------- | ---------- | --------------- |
| τ-bench airline (50 tasks)  | **0.887**   | 0.833      | 0.460           |
| τ-bench retail (114 tasks)  | **0.933**   | 0.886      | 0.623           |
| τ³-bench banking (97 tasks) | **0.313**   | 0.306      | 0.041           |

span-router passed 82.0% of airline tasks in all three trials, versus 72.0% for Jev Router. On retail, that was 86.8% versus 78.9%. Average airline task time was 28.6 seconds versus 32.9 seconds. The strongest gains are in airline and retail; banking is close.

span-router and Jev Router scores are averaged over three trials; OpenRouter Auto was run once. Every router used the same τ user simulator, GPT-4.1 at temperature 0. Measured September 26–27, 2026.

## Call it

Use the Respan Chat Completions endpoint. Change the model name, put your agent's context in the system message, and keep a stable thread identifier:

```python
response = client.chat.completions.create(
    model="span-router",
    messages=[
        {"role": "system", "content": "You are a support agent for an airline. Use the booking tools; never change a fare without confirmation."},
        *conversation_history,
    ],
    extra_body={"thread_identifier": "conversation_123"},
)
```

* Streaming, tool calls, and structured outputs are supported.
* Send the full message history with the same `thread_identifier` on every turn.
* Don't combine `span-router` with fallback or load-balancing options.
* Calls use Respan credits. Your own provider keys aren't supported.

## Pricing

Each request is priced using the model that served it, with no router markup:

* Uncached input and output use that model's rates.
* Cached input uses the model's cache-read rate.
* When a provider charges for cache writes, those tokens use its cache-write rate.

Check each request's cost in its Respan log.