Practical LLM Evals is a collection of model visualizers, developer utilities for testing inference, and more.
Repo: sujee/practical-llm-evals
Why I built it
Most LLM benchmarks answer broad questions about model quality. In day-to-day development, I often need a faster and more practical answer: how does this model behave on the endpoint I am actually using?
I built Practical LLM Evals to make those comparisons lightweight and repeatable. The tools run in the browser, work with OpenAI-compatible endpoints, and focus on the measurements developers care about when choosing a model or inference provider.
Tools
ZebraBench (graduated)
ZebraBench, a browser-based benchmark for LLM inference endpoints, started here as Quick LLM Bench. It has since graduated into its own project.
Model Visualizer
An interactive view of models available through Nebius Token Factory. It makes it easier to compare model intelligence, pricing, release timelines, and context-window trends—and to see the tradeoffs between capability and cost.
Explore the Model Visualizer → View the source →

Developer utilities and experiments
The repository also includes small utilities for smoke-testing Chat Completions and Responses APIs, measuring streaming throughput, and checking endpoint behavior.
I also use it as a home for experiential evaluations: playable games, simulations, and visual demos built with AI coding agents and open models. These projects reveal qualities—such as instruction following, tool use, iteration, and visual coherence—that are difficult to capture in a single benchmark score.
- Quake game built with coding agents and Kimi K3 →
- Kimi K3 demos and experiments →
- Developer testing utilities →
What I’m exploring
- Time to first token (TTFT): how quickly a model begins responding
- Decode speed: how fast it streams generated tokens
- End-to-end latency: how long a complete response takes
- Inference cost: how model pricing changes the practical cost of a task
- Model comparisons: speed, intelligence, context windows, and release trends
- Reasoning quality: how reliably models solve structured thinking tasks
- Coding and agent behavior: what models produce when paired with real coding agents and iterative workflows
This is intentionally a practical toolkit rather than a comprehensive benchmarking framework. The goal is to make useful comparisons quickly, inspect the behavior behind the numbers, and turn the results into better model and infrastructure choices.