New

  • Response formats: added a dedicated place to create, edit, and reuse saved response formats across model configurations.
  • Online evaluations: added alerts for online evaluation results.
product

Improved

  • Experiments: stabilized trace loading with clearer loading states, early row previews, and protection against stale results.
  • Experiment inference: reasoning models can now use reasoning and tool calls together during a run.
  • Models: updated model availability, pricing, context limits, recommended ordering, and BYOK filtering across the Platform and Enterprise catalogs.
  • Behaviors: improved custom behavior creation with better candidate selection and related reliability fixes.
  • Monitors and reports: standardized time, number, and unit formatting for clearer and more consistent results.
  • Agent: restored page context and skills, and corrected metric filter definitions.
  • Gateway: responses now preserve the exact model name provided in the original request.
  • Slack integration: improved channel caching and request handling to prevent rate-limit errors during setup.
product

Fixed

  • Log privacy: requests with logging disabled no longer retain request or response content.
  • Log filters: fixed list-based filters causing server errors and prevented invalid filters from silently broadening results.
  • Log exports: large exports now complete more reliably, accurately report delivered rows, and fail clearly instead of returning incomplete files.
  • Datasets: fixed trace imports and large row-removal operations failing on high-traffic accounts.
  • SQL editor: prevented non-admin users from accessing internal system columns and metadata while preserving supported log filtering.
  • Gateway: fixed empty provider chunks ending streams early and ensured failed streamed responses still produce logs.
  • Gemini usage: fixed responses incorrectly reporting zero token usage or cost.
  • Monitors: fixed high-volume rolling monitors missing scheduled evaluations or firing late.
  • OAuth: fixed organization switching and authentication issues in MCP integrations.
product

New

  • Gateway: added a native OpenRouter chat completions endpoint, with reasoning and provider preferences passed through correctly.
  • Agent search: Respan Agent can now use full-text search to find relevant information.
product

Improved

  • Behaviors: added LLM judging to improve behavior classification and analysis.
  • Experiments: experiment details now present grader names, outcomes, scores, latency, cost, tokens, and errors more clearly while hiding unnecessary metadata.
  • Experiment performance: added early row previews, lighter list responses, and compression for large experiment APIs so results appear sooner and transfer less data.
  • Model discovery: improved recommended-model ordering, BYOK availability filtering, and model-detail layouts to make suitable models easier to find and compare.
  • Model catalog: compressed model catalog responses to reduce loading time.
  • Workflows: optimized workflow execution paths for substantially faster processing.
  • Monitors: simplified monitor rendering for clearer, more consistent notifications.
product

Fixed

  • Log exports: large exports now use stable ordering, recover from storage interruptions, and avoid silently dropping rows. Failed or incomplete exports also report accurate progress and no longer provide untrustworthy downloads.
  • Reliability: brief cache interruptions no longer restart otherwise healthy web servers or turn into avoidable request failures.
  • Models: plain and provider-prefixed model names now resolve to the same catalog entry and behave consistently.
  • Spans: model icons remain readable when the Model column is narrow.
  • Lists: invalid sorting options now fall back safely instead of causing server errors.
  • Dashboard metrics: fixed organization-filtered quantile requests that could return server errors.
product

New

  • Evaluator results: LLM graders now provide a written rationale alongside each score, making it easier to understand why an evaluation passed or failed.
product

Improved

  • Model catalog: updated available and retired models, refreshed provider pricing, and expanded model capability details.
  • Model catalog loading: models now load on demand for each organization, with retry handling and protection against stale responses.
  • Log ingestion: very large log payloads are now captured reliably without stalling the ingestion pipeline.
  • Search: improved full-text search performance with a new ClickHouse text index.
  • Request logs: queries that exceed memory or time limits now return a clear, actionable response instead of a raw database error.
product

Fixed

  • Workflows: fixed an amplification issue that could cause workflows to perform unnecessary repeated work.
product

New

  • Error Tracking: added an incident-first investigation experience with impact summaries, timelines, contributing error groups, occurrence evidence, acknowledgements, and notes.
  • Spans: the “Aggregate into” control is now available to all Spans users, making it easier to investigate activity by supported groups.
  • Reports: reports now include Pulse metrics, and existing customer reports have been migrated to the updated format.
  • Agent chat: agent responses can now render charts, tables, and lists for clearer, more structured results.
  • Model catalog: added BYOK-only billing labels across the catalog and model selectors.
product

Improved

  • Experiments: refined evaluation analytics and comparison views while reducing unnecessary trace loading, and added backend status visibility and automatic refreshes for active experiments.
  • Evaluator editor: added expanded editors for LLM prompts and Python code, with immediate synchronization and smoother keyboard interactions.
  • Error Tracking: improved how errors are classified into incidents for more accurate investigation.
  • Model discovery: models now default to newest-first ordering, with clearer release dates and lifecycle information.
  • Model pages: improved layouts across compact, expanded, sparse, and data-rich views.
  • Experiment performance: compressed large experiment-log responses to reduce loading time and transferred data.
  • Hosted MCP: improved sign-in continuity for hosted MCP connections with a dedicated OAuth session flow.
  • Logs API: removed internal platform fields from log, trace, and span API responses.
product

Fixed

  • Experiments: fixed evaluator scoring so model, cost, and token metrics come from the experiment run instead of the source dataset row.
product

New

  • Microsoft Teams: Microsoft Teams is now available as a notification integration.
  • Reports: reports can now include Pulse summaries for richer incident and behavior reporting.
  • Logs: request activity can now be grouped by Threads or Traces directly from Logs, with filtering, sorting, refreshing, and detail panels available in each grouped view.
  • Experiments: experiment lists now include evaluator scores and model details, making results easier to compare without opening each experiment.
  • Behaviors: Behaviors is now available to regular platform users, with a complete draft, training, and review workflow.
  • Home: added a recently visited section for quickly returning to prompts and datasets.
product

Improved

  • Red Team: improved campaign setup, engine performance, CLI presentation, documentation, and error handling.
  • Experiments: conditional evaluator results can now return N/A. Aggregate scores exclude unscored rows and show the number of scored results, while histograms display N/A separately.
  • Evaluators: refined the grader editor, expanded Code evaluation access, and improved how LLM and Code configurations are displayed and saved.
  • Behaviors: improved classification accuracy, reduced false positives, added controls for flagging incorrect matches, and made behavior charts clearer across time ranges.
  • Error tracking: improved error classification and added automatic clustering for errors occurring around the same time.
  • Model catalog: improved model information and catalog performance, including richer descriptions, release dates, and coding benchmark scores.
  • Request logs: customer names and email addresses are now preserved in request-log lists, and nearby-request navigation more accurately follows the selected event’s timestamp.
  • Dashboard sessions: bookmarked dashboards now refresh expired sessions automatically instead of unexpectedly returning users to sign-in.
  • Metrics: empty breakdown charts now remain in a stable no-data state without flashing.
  • Prompts: new drafts are created only after an edit, keeping prompt version history cleaner.
  • Dataset imports: imported message histories now appear as conversations, and expected outputs are stored and displayed as formatted JSON.
product

Fixed

  • Experiments: boolean evaluator results now display false values correctly, and multi-turn prompt experiments retain the complete conversation history.
  • Experiment workflows: added task validation and made configured evaluator execution more consistent.
  • Prompts: fixed commit detection and redeployment behavior across all behavior-affecting settings. Load-balancing configuration is now preserved during JSON editing, and dropped parameters are persisted so they no longer reappear.
  • Evaluators: fixed commit messages overwriting evaluator descriptions and ensured each version displays its own commit message.
  • Security: closed an unauthorized live-event stream vulnerability and resolved an additional high-severity security issue.
product

New

  • Red Team: Red Team is now available online, with tools to create and run assessments, review results, and generate reports. The open-source Red Team engine is also available on PyPI.
  • Experiments: added filters for datasets, prompts, evaluators, and models, along with model, prompt, and evaluator details directly in experiment lists.
  • Dataset automation: automatic insert workflows now support trace datasets.
  • Dashboard autosave: named dashboard views now save automatically, making it easier to preserve dashboard changes as you work.
product

Improved

  • Experiments: added starring, provider icons for models, and server-side table sorting for more useful experiment management at scale.
  • Dashboards: improved synced chart-hover markers and comparison labels so tooltips and date ticks consistently respect the selected timezone.
  • Home dashboard: metrics refresh automatically in the background and update more reliably when returning to the page or changing the time range.
  • Metrics: charts now support copy-to-clipboard.
  • Behaviors: improved behavior detection accuracy and reduced false positives.
product

Fixed

  • Traces: span-level filters are now preserved when viewing grouped traces and threads, so results remain accurate.
  • Security: strengthened protections.
product

New

product

Improved

  • Dashboard charts: chart cards now stay above neighboring tiles when hovered or keyboard-focused, keeping chart controls accessible in dense dashboards.
  • Command palette: aliases now resolve during search, and the favorites action is now labeled “Pin.”
  • Onboarding: added a sign-out option during onboarding so users can switch accounts without getting stuck.
product