Smart AI Gateway prod · eu-central

One endpoint in front of six LLM providers. Every app calls the gateway; the gateway picks the fastest healthy provider, retries, and fails over automatically when one goes down. This console is what the operator sees.

Try it
Requests today
187,340
rpm now 312
Median latency
241ms
p99 1.9s
Cost per 1k tokens
$0.0007
riding free tiers · 6 keys, 1 place
30-day uptime
99.97%
providers healthy 6/6 · errors 0.02%

Provider fleet

Requests cascade through this chain. Kill one with “Simulate outage” and watch traffic re-route without a single dropped request.

Groq
p50 96ms · p99 230ms
operational
OpenRouter
p50 310ms · p99 744ms
operational
Cerebras
p50 88ms · p99 211ms
operational
Gemini
p50 420ms · p99 1008ms
operational
Cohere
p50 380ms · p99 912ms
operational
Cloudflare
p50 240ms · p99 576ms
operational

Requests / minute live

Green: primary routes. Cyan: re-routed traffic during a failover.

414 rpm0
primary routefailover re-route

Request log

Every call from every app in the fleet, as it happens. New rows stream in live.

timeroutemodelprovidertokenslatencystatus
14:41:59.87/v1/chatqwen-3-32bGroq1,377118ms200
14:41:55.29/v1/embeddingsmixtral-9x22Gemini1,946357ms200
14:41:51.28/v1/chatqwen-3-32bCerebras2,58191ms200
14:40:47.51/v1/audioqwen-3-32bOpenRouter2,114412ms200
14:40:43.54/v1/chatllama-4-70bOpenRouter3,618335ms200
14:40:39.85/v1/chatllama-4-70bCloudflare1,939239ms200
14:39:35.31/v1/chatqwen-3-32bOpenRouter2,083396ms200
14:39:31.82/v1/embeddingsglm-5-airCerebras2,29190ms200
14:39:27.24/v1/embeddingsglm-5-airCerebras2,18099ms200
14:38:23.22/v1/chatglm-5-airGemini1,377391ms200
14:38:19.63/v1/chatmixtral-9x22Cerebras1,234117ms200
14:38:15.23/v1/embeddingsmixtral-9x22Cloudflare2,531230ms200
14:37:11.72/v1/chatglm-5-airCloudflare1,245207ms200
14:37:07.37/v1/chatqwen-3-32bGemini493557ms200