ReliovaOpen-model token supply
Open-model tokens · metered inference

Open-model tokens, priced for production.

Buy prepaid access to DeepSeek, GLM and Kimi through an OpenAI-compatible API. Start with a small balance, observe real usage and performance, then top up as traffic grows.

Pay per tokenInput, cached input and output metered separately
Prepaid balanceStart small and top up when you are ready
OpenAI-compatibleChange the base URL, keep your existing SDK
Usage recordsTrack consumption by key and model
From test to production

Choose. Test. Top up.

The buying path is deliberately simple. Select the model family, run your workload with a small prepaid balance, and expand usage only after the route works for your application.

01 · CHOOSE

Match the model to the job

Compare model capability, context needs, thinking behavior and expected token mix. The calculator turns your workload into an estimated monthly balance.

Start hereModel shortlist and public selling price.
02 · TEST

Fund a small test balance

Use the same API shape your application already knows. Observe latency, output behavior, errors and billed usage on your own prompts before committing more traffic.

You receiveAPI key, endpoint details and usage visibility.
03 · SCALE

Top up as traffic grows

Move from evaluation to steady token consumption without changing the integration. Larger, predictable usage can be priced separately after the traffic shape is known.

Production pathPrepaid balance, metered drawdown and support.
Token supply

One balance. Multiple model routes. Continuous usage.

Your application sends an OpenAI-compatible request. Reliova routes it to the selected model, records billed token categories, and draws the cost from your prepaid balance.

DEEPSEEKGLMKIMIOTHERMETER + BALANCEYOUR APPLICATION
REQUESTMODELMETEROUTPUT
What changes the bill

“Cheap tokens” are not the same as a cheaper system.

A low list price can disappear after reasoning output, cache misses, retries, slow responses, and engineering overhead. We calculate the total operating picture around the behavior your product needs.

  • Real input/output distribution and multi-turn growth
  • Cache hit assumptions separated from cache misses
  • Thinking on/off behavior and billed reasoning tokens
  • Peak/off-peak or batch pricing windows
  • Failure, retry, and latency cost at expected concurrency
Open the cost calculator
Token catalog

Current public rates, with the billing categories visible.

Displayed prices come directly from Reliova’s live public catalog. Model availability and larger-volume pricing are confirmed before you fund a balance.

GLM-5.3

glm-5.3
$0.15Input / 1M
Cached / 1M
$0.52Output / 1M

Reliova public rate; cached input is billed as standard input unless stated otherwise.

GLM-5.2

glm-5.2
$0.15Input / 1M
Cached / 1M
$0.52Output / 1M

Reliova public rate; cached input is billed as standard input unless stated otherwise.

Kimi-K3

kimi-k3
$2.98Input / 1M
$0.30Cached / 1M
$14.90Output / 1M

Reliova public rate for the listed token categories.

DeepSeek V4 Pro

deepseek-v4-pro
$0.60Input / 1M
$0.022Cached / 1M
$2.01Output / 1M

Reliova public rate; thinking configuration can materially change billed output.

DeepSeek V4 Flash

deepseek-v4-flash
$0.22Input / 1M
$0.0070Cached / 1M
$0.67Output / 1M

Reliova public rate; optimized for cost-sensitive, latency-sensitive traffic.

Reliova published rate · Updated Sep 15, 2026 · 18:43 UTC. Loaded directly from Reliova’s live price API. Taxes and exceptional support are excluded.

Buy tokens

Ask for the current rate and a test balance.

Tell us the model, approximate monthly token volume and workload type. We will reply with the current public selling rate and the smallest practical test setup. Do not submit credentials, regulated data or proprietary prompts through this form.

Email: randy.qin@reliova.com

Start small

Fund a test balance. Measure real usage.

Tell us the model and expected token volume. We will return the current public selling rate and the smallest practical starting balance.