> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://respan.ai/docs/documentation/features/gateway/span-router/quickstart/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://respan.ai/_mcp/server. # Quickstart span-router is a router, not a model. Call it like any other model on Respan, and Respan picks the best model for each request. For how it selects models and the benchmark results, see [Concept](/docs/documentation/features/gateway/span-router/concept). ## Get started #### Check credits and create a key span-router runs on Respan credits or your organization's free allowance. It does not use your own provider keys. Make sure your organization has an available credit balance or allowance. Create a key under [API keys](https://platform.respan.ai/platform/api/api-keys), then export it: ```bash export RESPAN_API_KEY="YOUR_RESPAN_API_KEY" ``` For the Python example, install the OpenAI client: ```bash pip install openai ``` #### Send a Chat Completions request Use `https://api.respan.ai/api` as the base URL and `span-router` as the API model ID for span-router. Router requests are supported on `/api/chat/completions` only. **`cURL`** ```bash cURL curl -i https://api.respan.ai/api/chat/completions \ -H "Authorization: Bearer $RESPAN_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "span-router", "thread_identifier": "conv-8f2c", "messages": [ {"role": "system", "content": "You are a support agent for Acme."}, {"role": "user", "content": "My order #1234 never arrived."} ] }' ``` **`Python`** ```python Python import os from openai import OpenAI client = OpenAI( base_url="https://api.respan.ai/api", api_key=os.environ["RESPAN_API_KEY"], ) raw_response = client.chat.completions.with_raw_response.create( model="span-router", messages=[ {"role": "system", "content": "You are a support agent for Acme."}, {"role": "user", "content": "My order #1234 never arrived."}, ], extra_body={"thread_identifier": "conv-8f2c"}, ) response = raw_response.parse() print(response.choices[0].message.content) print(raw_response.headers.get("x-respan-log-id")) ``` Replace `conv-8f2c` with an identifier for your conversation. Keep that identifier for subsequent turns, and use a different one for each new conversation. #### Inspect the response and cost A successful request returns a normal Chat Completions response with `model` set to `span-router`. The response includes an `X-Respan-Log-Id` header. Keep this ID when investigating an unexpected result or contacting support. The request's log in the [Respan dashboard](https://platform.respan.ai) records its token usage and cost. The public model name remains `span-router`, even though an underlying model generated the answer. ## Conversations and prompt caching Send the full conversation history on every turn, including assistant messages and tool results. Chat Completions does not retain that history for you. Keep the same `thread_identifier` for each turn of a conversation. See [Keep continuity across turns](/docs/documentation/features/gateway/span-router/concept#keep-continuity-across-turns) and [Cache-aware cost](/docs/documentation/features/gateway/span-router/concept#cache-aware-cost). ## Supported features | Feature | Usage | | ----------------- | ------------------------------------------------------------------------------- | | Tools | Define `tools`. Use `tool_choice` to require a call or select a named function. | | Structured output | Request JSON mode or a JSON schema with `response_format`. | | Streaming | Set `stream: true`. Chunks use the `span-router` model name. | | Output limit | Set `max_tokens`. The limit includes reasoning tokens. | | Images | Include `image_url` content in a message. | Return tool results in the next request, validate structured responses against your application's schema, and test the image sizes and tasks your application uses. Reasoning tokens can consume the output limit before visible text is produced. ## Not supported span-router chooses the model and its fallbacks and always runs on Respan credits. These return an error: | You send | You get | | ----------------------------------------------------- | ----------------------------- | | `fallback_models`, `models` or load-balancing options | `400`: incompatible parameter | | Your own provider credentials | `400`: incompatible parameter | | Organization without credits | `402` | | Any endpoint other than Chat Completions | `400` | For example, sending `fallback_models` returns a `400` error that names `fallback_models` as the incompatible parameter. If a request fails or returns an unexpected answer, send its `X-Respan-Log-Id` and what you expected to [Respan support](/docs/documentation/support). ## Pricing Router requests use Respan credits or your organization's free allowance. See [Pricing](/docs/documentation/features/gateway/span-router/concept#pricing) for model rates and cache billing. ## Benchmarks See the [agent-task benchmarks](/docs/documentation/features/gateway/span-router/concept#agent-tasks) for task completion rates. See the [single-request benchmarks](/docs/documentation/features/gateway/span-router/concept#single-requests) for quality and cost comparisons. > Send your first request to span-router.