Each model is probed with a synthetic request every 20 minutes (Cloudflare cron) · monitoring freeinference.org . Click a model to zoom in on its latency and throughput history.
bge-m3
Status UP
Latency 145 ms
TTFT —
Throughput —
Uptime 99%
Checked 2026-10-11T09:41:06
Latency trend (99) 145 ms · 127–339 ms
click to zoom ↗
deepseek-v4-flash
Status UP
Latency 3189 ms
TTFT 257 ms
Throughput 234.31 tok/s
Uptime 99%
Checked 2026-10-11T09:41:02
TTFT trend (99) 257 ms · 221–12520 ms
click to zoom ↗
diffusiongemma
Status UP
Latency 4687 ms
TTFT 3638 ms
Throughput 218.48 tok/s
Uptime 99%
Checked 2026-10-11T09:40:58
TTFT trend (99) 3638 ms · 2634–7599 ms
click to zoom ↗
glm-5.3
Status UP
Latency 14960 ms
TTFT 2552 ms
Throughput 82.45 tok/s
Uptime 100%
Checked 2026-10-11T09:40:22
TTFT trend (99) 2552 ms · 1875–23324 ms
click to zoom ↗
glm-5.3-flash
Status UP
Latency 2362 ms
TTFT 471 ms
Throughput 145.95 tok/s
Uptime 100%
Checked 2026-10-11T09:40:37
TTFT trend (100) 471 ms · 192–11402 ms
click to zoom ↗
kimi-k2.7-code
Status DOWN
Latency 2087 ms
TTFT —
Throughput —
Uptime 78%
Checked 2026-10-11T09:40:55
TTFT trend (78) 1441 ms · 1235–2833 ms
click to zoom ↗
HTTP 403: You've reached your concurrent request limit. Please wait for your ongoing requests to finish and try again. (request_id: req_6e3728cbe74b446f850f2614b0c183d1)
minimax-m2.5
Status UP
Latency 8832 ms
TTFT 3225 ms
Throughput 137.15 tok/s
Uptime 92%
Checked 2026-10-11T09:40:39
TTFT trend (92) 3225 ms · 1095–7509 ms
click to zoom ↗
minimax-m3
Status UP
Latency 6643 ms
TTFT 780 ms
Throughput 102.85 tok/s
Uptime 100%
Checked 2026-10-11T09:40:48
TTFT trend (100) 780 ms · 625–45151 ms
click to zoom ↗
qwen3.6-35b
Status UP
Latency 1054 ms
TTFT 319 ms
Throughput 342.86 tok/s
Uptime 100%
Checked 2026-10-11T09:40:57
TTFT trend (100) 319 ms · 140–412 ms
click to zoom ↗