Developers

Status

Live availability for every component, with 90 days of history and a full write-up for every incident.

All systems operational Last checked 30 seconds ago · updates every 30s
API gateway global Operational 99.99%
90 days agotoday
Router global Operational 99.99%
90 days agotoday
Model pool — frontier multi-region Operational 99.96%
90 days agotoday
Model pool — open weights multi-region Operational 99.99%
90 days agotoday
Observability ingest global Operational 99.98%
90 days agotoday
Dashboard global Operational 100.00%
90 days agotoday
Webhooks global Operational 99.97%
90 days agotoday

Availability is measured from synthetic probes in every region, not from self-reported health. A day is marked degraded if error rate exceeded 0.1% or p95 latency exceeded twice the trailing median.

Uptime

Rolling averages across all services.

  • 99.993%Last 30 days
  • 99.991%Last 90 days
  • 99.99%Trailing 12 months
  • 3Incidents in 90 days
Incident history

Every incident, with the timeline.

Elevated latency on frontier model pool (us-east)

Resolved 2026-07-14
  • Investigating — p95 latency on us-east frontier routes rose above 6s. Router began shifting traffic to us-west.
  • Identified — an upstream provider was returning 529 under load. Affected models were marked degraded and removed from rotation.
  • Monitoring — latency back within normal range. Requests continued to succeed throughout via fallback.
  • Resolved — provider capacity restored. No requests were dropped; 0.4% of requests took a second attempt.

Delayed trace ingestion

Resolved 2026-06-08
  • Investigating — traces were appearing in the dashboard with up to 9 minutes of delay.
  • Identified — a backlog in the ingest queue after a bad deploy of the span writer. Inference was unaffected.
  • Monitoring — the deploy was rolled back and the backlog drained.
  • Resolved — trace availability back under 5 seconds. No telemetry was lost; spans were buffered and replayed.

Router degradation in eu-central

Resolved 2026-05-19
  • Investigating — a subset of eu-central requests returned 500 from the router.
  • Identified — a config push shipped an invalid fallback chain for one project tier.
  • Resolved — config reverted. 0.02% of eu-central requests errored for 21 minutes. Affected accounts were credited.

Incidents older than 90 days are kept in the archive and are available to Enterprise accounts on request.