ReliovaOpen-model token supply
Reproducible by design

Evaluation methodology

We treat model, serving route, configuration, and workload as separate variables. A conclusion applies only to the combination actually tested.

1. Freeze the question

We document the business task, acceptance criteria, traffic shape, latency target, data restrictions, current baseline, and failure cost. Before testing, we agree which outcomes would justify a move and which would stop it.

2. Build a representative set

Customer-provided examples are sampled across normal, difficult, long-context, multilingual, and tool-use cases. Sensitive content is minimized or replaced where possible. The customer approves the dataset and scoring method.

3. Control the run

Every candidate receives equivalent messages, tools, output constraints, sampling settings, and retry policy where the APIs permit it. Model identifiers, timestamps, endpoint route, response headers, usage fields, and errors are logged. Unsupported settings are reported rather than silently normalized.

4. Measure separate axes

Quality

Deterministic checks, rubric scoring, human review, and explicit failure categories.

Performance

Time to first token, generation rate, end-to-end latency, error rate, and behavior under bounded concurrency.

Economics

Fresh input, cached input, visible output, reasoning output, retries, and price window.

Operations

Rate limits, support path, data handling, change control, billing evidence, and token-supply fit.

5. Report uncertainty

Results include sample size, test dates, observed configuration, confidence limits where useful, and known gaps. We do not infer hidden model weights or quantization from a magic token count. Output differences can be evidence of a serving difference, but they are not proof of its cause without controlled follow-up.

6. Re-test before production

A short benchmark is not a production guarantee. Readiness requires a capacity window, failure injection, billing reconciliation, data-path review, and a rollback path. Changes in model version or route invalidate relevant parts of the earlier result.

Start small

Fund a test balance. Measure real usage.

Tell us the model and expected token volume. We will return the current public selling rate and the smallest practical starting balance.