Investigate error incidents

Find error spikes, identify contributing failures, and coordinate the response.

The Errors page groups failed spans into incidents and error groups. Use it to measure impact, find the failures behind an outage, and inspect matching events.

If Errors does not appear in the navigation, contact your Respan administrator or Respan support.

What you can answer

  • When did an error spike start, and is it still ongoing?
  • Which providers, models, endpoints, or customers were affected?
  • Which error groups contributed most?
  • Has someone acknowledged and investigated the incident?

Read the error timeline

Errors uses two levels of grouping:

LevelMeaning
IncidentA window where one or more error groups spike or drift above their baseline.
Error groupFailures with the same error class, provider, endpoint, and HTTP status.

The streamgraph shows error volume across the selected range. Colored layers represent the ten largest error groups; the rest are combined into Other errors. Hover a legend item or group row to isolate its layer.

Incident rows show acknowledgement, state, severity, error count, and start time. Hover a row to highlight its window, then select it for details. Error-group rows show an aligned occurrence trend and Unresolved, Resolved, or Ignored state.

Incidents are detected from error-rate behavior. They are separate from monitors configured by your organization.

Set the scope

1

Choose the environment and time range

Include the incident and enough time before it to show the normal baseline.

2

Apply filters

Filter by error type, provider, fault domain, or HTTP status. Fault domain affects incidents and groups; the other filters narrow the groups on the timeline.

3

Save the view

Use All incidents to save a filter set that your team needs to revisit.

Investigate an incident

Select an incident to open its detail panel. The header summarizes its error rate, environment, duration, affected scope, and change from the normal rate. Tags show severity, dominant error type, state, acknowledgement, and whether impact is estimated.

Use the detail sections to:

  • Rank Contributing error groups by share and occurrence count. Select one to drill into it while retaining a breadcrumb to the incident.
  • Review affected providers, models, endpoints, and customers.
  • Compare failed and total requests, normal and peak error rates, start and end times, trigger, fault domain, and latest Data through timestamp.
  • Select Acknowledge when someone takes ownership.
  • Add investigation context under Internal notes.

Internal notes are shared. Do not include secrets, credentials, customer content, or raw prompts.

Ongoing incidents refresh every minute while the page is visible. Check Data through before treating current counts as final.

Investigate an error group

Open a group from the timeline or an incident’s contributor list. A contributor drill-down scopes headline counts to the incident; sampled events continue to use the page’s time range.

The Overview tab shows occurrences, diagnostics, fingerprint metadata, and a representative event payload. Use the status control to Resolve, Ignore, or Reopen the group.

Open Events to inspect sampled occurrences. Break them down by Model, Customer, or Environment, select values to narrow the list, then select an event for its log detail.

Events are sampled, not an exhaustive export. Use the occurrence total for impact and Logs for broader record-level analysis.

Triage an incident

  1. Choose the affected environment and a range with a short baseline.
  2. Check the incident severity, error rate, duration, and Data through value.
  3. Review contributing groups and affected dimensions.
  4. Acknowledge the incident and add an ownership note.
  5. Open the dominant group and inspect its diagnostics and events.
  6. Continue in Logs or Traces to confirm the first failing span and root cause.
  7. Resolve handled groups and create a monitor for a measurable recurring condition.

Choose the right view

UseWhen you need to
ErrorsTriage incidents and grouped failures.
MetricsCompare error trends with traffic, latency, and cost.
LogsInspect individual failed spans.
TracesUnderstand the execution around a failed span.
MonitorsNotify a destination when a metric crosses a threshold.

Troubleshooting

The timeline has no incidents

Widen the range, clear filters, and confirm the environment. Error groups can appear without crossing the threshold for an incident; check Logs if both are empty.

Filtering hides unexpected data

Fault domain applies to incidents and groups. Error type, provider, and HTTP status affect the groups, not the incident list. Clear filters, then add them one at a time.

An expected event is missing

The Events tab is sampled. Open Logs with the same environment, time range, fingerprint, provider, and HTTP status, then search for the event or request ID.


Need help?

Join our Discord — we’ll help you investigate an error incident.