ZebraBench is a quick, browser-based benchmark for testing one or more LLM inference endpoints—without spinning up a VM, setting up a Python environment, or installing packages.
Point it at any OpenAI-compatible endpoint, enter an API key, load the available models, and select the models you want to compare. It measures time to first token, token throughput, end-to-end latency, accuracy, token usage, and estimated cost. Results can be exported as CSV or JSON for further analysis.
Repo: sujee/zebrabench
Try ZebraBench → View the source →

What it measures
- Time to first token (TTFT): how quickly a model begins responding
- Decode speed: how fast it streams generated tokens
- End-to-end latency: how long a complete response takes
- Accuracy: how reliably a model answers structured tasks
- Token usage and cost: how pricing changes the practical cost of a task
This is designed to be a quick benchmark. It is not intended to replace comprehensive benchmarking tools.
ZebraBench started as part of my Practical LLM Evals toolkit and has since graduated into its own project.