Decision summary
Conditional go: Candidate B meets the quality floor and lowers modeled unit cost, but should not receive interactive traffic until p90 TTFT and the error behavior at target concurrency are accepted in a production qualification window.
Test definition
| Workload | Customer support summarization + structured extraction |
|---|---|
| Sample | 500 de-identified requests, stratified by length and difficulty |
| Quality floor | ≥ baseline pass rate; zero new critical extraction errors |
| Traffic target | Illustrative only: 20 concurrent requests |
| Run window | Recorded with UTC timestamps and model identifiers |
What the report shows
- Pass rate and failure taxonomy with reviewed examples
- p50, p90, and p99 latency—not a single average
- Output tokens per second after first token
- Error and retry counts at each concurrency level
- Usage-field reconciliation and modeled cost range
- Data path, retention, support, and operational questions
- Migration gates, rollback trigger, and recommended next test
What it does not claim
The report does not certify an underlying model’s undisclosed weights, promise future performance, or generalize beyond the tested route and configuration. Production details are confirmed for the selected service configuration.