FreeInference Model Status

Each model is probed with a synthetic request every 20 minutes (Cloudflare cron) · monitoring freeinference.org. Click a model to zoom in on its latency and throughput history.
8 up 1 down 9 models
bge-m3
StatusUP Latency145 ms TTFT— Throughput— Uptime99% Checked2026-10-11T09:41:06
Latency trend (99)145 ms · 127–339 ms
click to zoom ↗
deepseek-v4-flash
StatusUP Latency3189 ms TTFT257 ms Throughput234.31 tok/s Uptime99% Checked2026-10-11T09:41:02
TTFT trend (99)257 ms · 221–12520 ms
click to zoom ↗
diffusiongemma
StatusUP Latency4687 ms TTFT3638 ms Throughput218.48 tok/s Uptime99% Checked2026-10-11T09:40:58
TTFT trend (99)3638 ms · 2634–7599 ms
click to zoom ↗
glm-5.3
StatusUP Latency14960 ms TTFT2552 ms Throughput82.45 tok/s Uptime100% Checked2026-10-11T09:40:22
TTFT trend (99)2552 ms · 1875–23324 ms
click to zoom ↗
glm-5.3-flash
StatusUP Latency2362 ms TTFT471 ms Throughput145.95 tok/s Uptime100% Checked2026-10-11T09:40:37
TTFT trend (100)471 ms · 192–11402 ms
click to zoom ↗
kimi-k2.7-code
StatusDOWN Latency2087 ms TTFT— Throughput— Uptime78% Checked2026-10-11T09:40:55
TTFT trend (78)1441 ms · 1235–2833 ms
click to zoom ↗
HTTP 403: You've reached your concurrent request limit. Please wait for your ongoing requests to finish and try again. (request_id: req_6e3728cbe74b446f850f2614b0c183d1)
minimax-m2.5
StatusUP Latency8832 ms TTFT3225 ms Throughput137.15 tok/s Uptime92% Checked2026-10-11T09:40:39
TTFT trend (92)3225 ms · 1095–7509 ms
click to zoom ↗
minimax-m3
StatusUP Latency6643 ms TTFT780 ms Throughput102.85 tok/s Uptime100% Checked2026-10-11T09:40:48
TTFT trend (100)780 ms · 625–45151 ms
click to zoom ↗
qwen3.6-35b
StatusUP Latency1054 ms TTFT319 ms Throughput342.86 tok/s Uptime100% Checked2026-10-11T09:40:57
TTFT trend (100)319 ms · 140–412 ms
click to zoom ↗