Skip to navigation

Quickstart

Send your first request to span-router.

span-router is a router, not a model. Call it like any other model on Respan, and Respan picks the best model for each request. For how it selects models and the benchmark results, see Concept.

Get started

1

Check credits and create a key

span-router runs on Respan credits or your organization’s free allowance. It does not use your own provider keys. Make sure your organization has an available credit balance or allowance.

Create a key under API keys, then export it:

export RESPAN_API_KEY="YOUR_RESPAN_API_KEY"

For the Python example, install the OpenAI client:

pip install openai
2

Send a Chat Completions request

Use https://api.respan.ai/api as the base URL and span-router as the API model ID for span-router. Router requests are supported on /api/chat/completions only.

curl -i https://api.respan.ai/api/chat/completions \
-H "Authorization: Bearer $RESPAN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "span-router",
"thread_identifier": "conv-8f2c",
"messages": [
{"role": "system", "content": "You are a support agent for Acme."},
{"role": "user", "content": "My order #1234 never arrived."}
]
}'

Replace conv-8f2c with an identifier for your conversation. Keep that identifier for subsequent turns, and use a different one for each new conversation.

3

Inspect the response and cost

A successful request returns a normal Chat Completions response with model set to span-router. The response includes an X-Respan-Log-Id header. Keep this ID when investigating an unexpected result or contacting support.

The request’s log in the Respan dashboard records its token usage and cost. The public model name remains span-router, even though an underlying model generated the answer.

Conversations and prompt caching

Send the full conversation history on every turn, including assistant messages and tool results. Chat Completions does not retain that history for you.

Keep the same thread_identifier for each turn of a conversation. See Keep continuity across turns and Cache-aware cost.

Supported features

FeatureUsage
ToolsDefine tools. Use tool_choice to require a call or select a named function.
Structured outputRequest JSON mode or a JSON schema with response_format.
StreamingSet stream: true. Chunks use the span-router model name.
Output limitSet max_tokens. The limit includes reasoning tokens.
ImagesInclude image_url content in a message.

Return tool results in the next request, validate structured responses against your application’s schema, and test the image sizes and tasks your application uses. Reasoning tokens can consume the output limit before visible text is produced.

Not supported

span-router chooses the model and its fallbacks and always runs on Respan credits. These return an error:

You sendYou get
fallback_models, models or load-balancing options400: incompatible parameter
Your own provider credentials400: incompatible parameter
Organization without credits402
Any endpoint other than Chat Completions400

For example, sending fallback_models returns a 400 error that names fallback_models as the incompatible parameter.

If a request fails or returns an unexpected answer, send its X-Respan-Log-Id and what you expected to Respan support.

Pricing

Router requests use Respan credits or your organization’s free allowance. See Pricing for model rates and cache billing.

Benchmarks

See the agent-task benchmarks for task completion rates.

See the single-request benchmarks for quality and cost comparisons.