Command Palette

Search for a command to run...

GitHubRun the benchmark
Benchmarks / AI Inference

AI Inference benchmarks

Time to first token, streaming throughput, cold starts and tail latency across AI inference providers — measured continuously, not quoted from launch blogs.

Planned — not yet measuring
What we'll measure
  • Time to first token
  • Tokens per second
  • p99 request latency
  • Cold start time
  • Price per 1M tokens
Providers on the roadmap
OpenAIAnthropicGoogleGroqTogether AIFireworks AI

Same rules as compute: deterministic workloads, open data, no affiliations. The test suite will ship in the same CLI — providerbench run -t ttft — so anyone can verify from their own network.

Want this category sooner?

The benchmark framework is one Go interface — category suites are designed in the open. Propose metrics or contribute tests.

Open a discussion