Arena Measurement Guide
Qorinix Arena compares responses to the same prompt across the available inference lanes. Each run is a useful observation, not an independent certification of model quality or a universal speed ranking.
Time to first token
TTFT describes how long you wait for the response to start. The public interface observes the response path, which can include browser, network, routing and upstream inference time. It is not an isolated measurement of model execution.
Completion time and throughput
Completion time also depends on output length. Compare prompts, output constraints and conditions consistently. Token estimates and derived throughput should be interpreted in the context of the displayed run.
Quality and reliability
Read the outputs and check them against the task requirements. A fast answer can be incorrect. Repeat tests and record failures, timeouts and quota limits instead of ignoring them.
Benchmark dashboard
The Benchmark dashboard uses a fixed illustrative dataset for exploring sorting, charts and score tradeoffs. The displayed quality scores, costs and timing figures are examples. It does not ingest a live rolling production dataset.
Reproducible evaluations
- Choose a representative prompt set and fixed output constraints.
- Run each prompt repeatedly under comparable conditions.
- Record first-response and completion times with errors.
- Review correctness separately from speed.
- Document the test date and scope when sharing results.