DeepSeek V4 Flash

DeepSeektext

Cheaper, faster V4 tier with a ~98% cache-hit discount; extremely economical for repeated prompts.

DeepSeek V4 Flash strengths

  • Among the cheapest capable models
  • Massive cache savings
  • 1M context
  • Good general performance

Pricing & context

Context window1M tokens (384K max output)
Input price /1M$0.14 ($0.0028 cache hit)
Output price /1M$0.28
Modalitiestext

Cost guide: a typical call of about 10K input + 2K output tokens costs roughly $0.002 at list prices. Worth modelling against cheaper tiers before committing high-volume traffic.

When to choose DeepSeek V4 Flash

DeepSeek V4 Flash is best for high-volume, repetitive workloads and budget pipelines needing long context. If your workload is more cost-sensitive, weigh it against gpt-oss-120b (≈$0.03 input /1M) first.

DeepSeek V4 Flash FAQ

How much does DeepSeek V4 Flash cost?

DeepSeek V4 Flash is priced at $0.14 ($0.0028 cache hit) per 1M input tokens and $0.28 per 1M output tokens (public API list price), with a 1M tokens (384K max output) context window. A typical call of about 10K input and 2K output tokens costs roughly $0.002.

What is DeepSeek V4 Flash best for?

DeepSeek V4 Flash by DeepSeek is best for high-volume, repetitive workloads and budget pipelines needing long context.

How does DeepSeek V4 Flash pricing compare to Doubao Seed 2.1 Turbo?

DeepSeek V4 Flash input costs $0.14 ($0.0028 cache hit) per 1M tokens versus $0.45 for Doubao Seed 2.1 Turbo, roughly 3.2x less expensive on input. Output is $0.28 vs $2.25.

Is DeepSeek V4 Flash multimodal?

DeepSeek V4 Flash supports text.

Tools that use DeepSeek V4 Flash

Other models

All models →
01Claude Fable 5Anthropic$10.00
02Claude Opus 5Anthropic$5.00
03GPT-5.6 SolOpenAI$5.00
04GPT-5.5OpenAI$5.00