English
English
With Claude the hard part is access: cards get declined, regions get restricted, and a relay is often the only practical route. DeepSeek is different. Its platform is reachable, its documentation publishes the rates, and its base URL speaks the OpenAI shape already. So the usual "who will sell me a key" question mostly does not apply.
What remains is an accounting question, and it has a shape. DeepSeek bills input and output separately, splits input into cache hit and cache miss, and applies different rates at peak and off-peak hours. Four numbers, not one. A single "per million tokens" headline collapses all four into something that cannot be compared.
How to check: open the official pricing page and count the columns. If a provider quotes you one number for DeepSeek, ask which of the four it covers. That question alone sorts the serious operators from the ones reading from a screenshot.
| Line item | What triggers it | What makes it grow | Who controls it |
|---|---|---|---|
| Input, cache miss | Any prompt prefix that is not already in the cache | Resending fresh context on every call | Your application |
| Input, cache hit | A prompt whose prefix exactly matches one already persisted | Reusing a stable prefix | Your application |
| Output | Every token the model generates | Long answers, verbose reasoning, retries | Your prompt and settings |
| Peak multiplier | Calls that land inside the provider's peak window | Batch jobs scheduled at the wrong hour | Your scheduler |
The gap between the first two rows is the largest single lever on the page. In DeepSeek's own table the cache-hit input rate is a small fraction of the cache-miss rate — the kind of difference that turns a cost estimate upside down.
The last row is the one almost nobody plans for. Because official rates move between peak and off-peak, a nightly batch and an afternoon batch of identical size are not the same purchase. If your workload is not latency-sensitive, that is free money.
This is the part worth reading the documentation for, because the rules are more specific than "repeated prompts are cheaper".
DeepSeek's context caching is enabled by default for all users and runs on disk. Every request triggers cache construction. A later request hits the cache only when its prefix fully matches a persisted cache prefix unit — a partial overlap does not qualify.
The units are created at request boundaries: one at the end of the user input and one at the end of the model output. That is why the practical advice is not "send the same prompt twice" but "keep an identical, stable block at the very front of every request in a session".
How to check: send the same long prefix twice and read the usage block on the second response. If cache-hit input tokens appear, caching is doing its job for your traffic. If they never appear, the prefix is changing somewhere — a timestamp, a re-ordered tool list, a rewritten system message will all do it.
Run these in order. Each step produces the evidence the next one needs, and the whole sequence costs less than one page of output.
1. Ask the endpoint which DeepSeek models it serves. Take the identifier from the live list, not from an article. Model identifiers on this family age faster than documentation does, and a stale identifier looks exactly like an outage.
2. Send one minimal completion. Small output limit, one-line prompt, OpenAI-compatible path. You are not reading the answer — you are reading the envelope.
3. Read the usage block. A compliant response separates input from output, and separates cache hits from cache misses when a cache was touched. A first short call usually shows no cache figures. That is expected.
4. Reconcile that one request in the usage record. Compare the token counts and the amount deducted against step 3. This is the only step where price and usage meet in a single row, which is why it is the step that separates an auditable service from a plausible one.
5. Force one failure, then look at the record. Send a call with a model identifier that cannot exist. On KuaiAPI failed requests are not billed — that is a published commitment, which makes it a claim you can test in a minute rather than a favour you hope for.
6. Repeat step 2 with a long, identical prefix. The second call is where cache-hit tokens should appear, priced differently from fresh input. This is the single most informative number you can get about a DeepSeek endpoint.
If step 6 shows nothing, you have learned something more useful than any rate comparison: caching is not reaching your traffic, and the headline number is not the number you will pay.
A relay states its DeepSeek price as the official rate multiplied by a group multiplier, and a good one publishes both halves. On KuaiAPI that decomposition is on the model plaza: each group shows its multiplier, and each model inside the group shows input, output and cache-read rates separately, refreshing from live data rather than from a hard-coded table.
That structure is what makes a claim checkable. You can multiply the official rate by the published multiplier yourself and see whether the result matches the rate shown. If a provider publishes only the result, there is nothing to check.
How to check: pick one DeepSeek model, read its group multiplier, and do the multiplication against the official page. Two minutes, and it either reconciles or it does not.
Keep the prefix stable across a session. Caching rewards an identical front; every framework that rewrites the system prompt per call quietly cancels the discount. Pin it, timestamp it outside the cached block, and the same traffic gets cheaper without changing models.
Schedule batch work off-peak. Peak and off-peak rates are published, so this is a scheduling decision rather than a negotiation. If a job can wait, letting it wait is the cheapest optimisation available.
Watch output tokens, not just input. Reasoning-heavy modes generate more than you asked for, and output is the expensive direction. An output limit is not a quality knob here; it is a billing one.
Is a DeepSeek API key free or paid? The API is paid; the chat product is a separate thing and is not the same offer. Treat any page that promises a free DeepSeek API key as a page about something else, and check the official pricing page for what is actually being sold.
Is OpenRouter DeepSeek free? This question is common enough to appear in Google's own related-question blocks, and the honest answer is that aggregation adds a routing layer, not a discount. What you get from an aggregator is one credential for several model families; whether that is worth more than the margin depends on how many families you actually call.
Why is my bill higher than the list price suggests? Almost always cache misses. If the front of your prompt changes on every call, every call is billed as fresh input, and the cache-hit column never applies to you. Check the usage block before changing providers — the problem usually travels with the application.
Do I need a different endpoint for DeepSeek than for GPT or Claude? DeepSeek's own platform publishes both an OpenAI-format and an Anthropic-format base URL, so both client families are served. On an aggregator, the practical rule is the same as for any other model: OpenAI-compatible clients point at the /v1 path, Claude Code points at the host, and the credential is the same.
The two pages that answer most remaining questions need no login. The model plaza lists each group's multiplier and each model's input, output and cache-read rates, so a price can be decomposed instead of compared. The status page shows how each group is behaving right now, including when it is not behaving well.
For the surrounding decisions, Cheap Claude API: How Per-Million-Token Prices Really Compare walks through the same four-line-item structure on a different model family, What Is an AI API Aggregator? covers the operator-level questions, and One API Key for All Models shows the client-side setup when one credential has to serve several tools. The English documentation covers endpoints and credential configuration.
Ready to start calling?
Create an account, generate a key and make your first multi-model call in five minutes. One key for GPT, Claude, Gemini, Grok, DeepSeek and Chinese models.