DeepSeek V4 Flash Pricing: What the 3¢ Benchmark Test Means

Home News DeepSeek V4 Flash Pricing: What the 3¢ Benchmark Test Means
AI & Automation

DeepSeek V4 Flash pricing is $0.14 input and $0.28 output per million tokens. Artificial Analysis puts benchmark-task cost near 3¢.

PK
August 4, 2026 5 min

Artificial Analysis has put a striking number on DeepSeek V4-Flash-0731: about $0.03 for the weighted average cost of one task in its Intelligence Index. The research firm also scored the model at 50, a 10-point gain over the V4-Flash version released in April. That makes the July 31 update notable on both capability and cost, but the result is narrower than a “three-cent AI task” headline suggests.

DeepSeek’s official API table lists $0.14 per million cache-miss input tokens, $0.28 per million output tokens, and $0.0028 per million cached input tokens. The same page says peak-hour rates will eventually be twice the regular prices, with the effective date still to be announced. Token rates describe what the API meters. Artificial Analysis’s task figure describes what one benchmark attempt consumed under its evaluation method.

For enterprise buyers, neither number answers the final budgeting question: what does a usable, approved outcome cost after retries, tool calls, fallbacks, and human review? Our read: V4-Flash has earned a place in an enterprise evaluation set, not a blanket migration plan.

Direct answer: What does DeepSeek V4-Flash’s three-cent benchmark cost mean?

DeepSeek V4 Flash pricing on the first-party API is $0.14 per million cache-miss input tokens, $0.28 per million output tokens, and $0.0028 per million cache-hit input tokens. Artificial Analysis estimates a weighted benchmark cost of about three cents per Intelligence Index task. That is cost per benchmark attempt under its test mix, not a guarantee that every successful production outcome costs three cents.

Key Takeaways

  • Artificial Analysis estimates about $0.03 per Intelligence Index task and gives V4-Flash-0731 a score of 50.
  • DeepSeek lists $0.14 per million cache-miss input tokens, $0.28 for output, and $0.0028 for cached input.
  • DeepSeek says future peak-hour API rates will be 2x regular prices; the effective date is not announced.
  • Artificial Analysis’s three-cent metric is cost per benchmark attempt, not cost per correct or accepted production outcome.
  • Enterprise teams should compare cost per accepted task, including retries, tools, fallbacks, and human review.

What DeepSeek V4 Flash Pricing Actually Includes

DeepSeek’s July 31 changelog says V4-Flash-0731 keeps the preview model’s architecture and size and was re-post-trained. The API name deepseek-v4-flash now routes to the updated model. Its first-party bill separates cache hits, cache misses, and output, so the effective blended rate depends on prompt reuse and response length.

Published unit pricing improves calculability, but it does not establish total cost or product quality. Our 100-vendor AI pricing source revalidation made the same boundary explicit: a buyer can know the marginal usage price and still lack the evidence needed to judge value. Provider markups, future peak pricing, and workload behavior can all change the invoice.

What the Three-Cent Benchmark Cost Measures

Artificial Analysis defines cost per Intelligence Index task as a weighted average. For each evaluation, it combines input, cache-hit, cache-write, reasoning, and answer-token prices with the tokens used, divides by task count, and applies the evaluation’s Index weight. That is why the metric is more informative than list price alone: it captures how many priced tokens the model used to attempt the work.

The denominator is task count, not correct or enterprise-approved outputs. For production, teams need a second calculation: cost per accepted task = total model, orchestration, tool, retry, and review cost divided by accepted outputs. A model can look cheap per token or per attempt and still become expensive when failures, long responses, or human review multiply.

Where V4-Flash Sits on Artificial Analysis’s Benchmark Curve

On Artificial Analysis’s Intelligence Index, V4-Flash-0731’s score of 50 matches Gemini 3.6 Flash, sits one point behind Muse Spark 1.1 and GLM-5.2, and seven points behind Kimi K3. In the firm’s cost-per-task comparison, V4-Flash was about $0.03, versus $0.86 for Kimi K3, $1.86 for GPT-5.6 Sol, and $3.15 for Claude Fable 5.

Artificial Analysis also places V4-Flash-0731 on its Pareto frontier for intelligence versus cost per task. Those are figures for its benchmark mix, tested settings, and selected providers, not a universal price-performance ranking. The firm lists V4-Flash-0731 as text-only, while competing models can differ in modality, latency, context behavior, and governance. The defensible conclusion is narrower: V4-Flash occupies a strong cost-intelligence point in this evaluation.

What Enterprise AI Teams Should Test Before Scaling

Define success before price. Build a representative evaluation set from real prompts, then set pass criteria for factuality, format, tool use, latency, and review. Do not let a composite benchmark score substitute for the failure modes that matter in your workflow.

Instrument the whole cost stack. Log cache hits, uncached input, reasoning, output, retries, tool calls, fallbacks, and human review. Re-run the test during any future peak-pricing window because DeepSeek says those rates will be twice the regular schedule.

Route by workload instead of replacing one model everywhere. V4-Flash may fit high-volume text tasks that meet your acceptance threshold, while higher-risk or low-confidence cases move to another model or a reviewer. That architecture turns low-cost inference into a controlled option rather than a single-vendor bet.

Check enterprise controls before scaling. Data handling, regional availability, service levels, version pinning, incident response, and provider support can outweigh a low inference bill. Cheaper accepted tasks can make more extraction, classification, summarization, and offline evaluation economically viable. The savings are only real after the production denominator is measured.

Frequently Asked Questions

DeepSeek lists V4-Flash-0731 at $0.14 per million cache-miss input tokens, $0.28 per million output tokens, and $0.0028 per million cache-hit input tokens on its first-party API. It also says future peak-hour rates will be twice the regular prices, with the start date not yet announced.

No. Artificial Analysis’s three-cent figure is a weighted average cost per Intelligence Index task attempt. Its denominator is benchmark task count, not correct answers or enterprise-approved outcomes. Production cost also includes retries, tool calls, fallbacks, orchestration, and human review, so teams should calculate cost per accepted task separately.

Artificial Analysis scored V4-Flash-0731 at 50. That matched Gemini 3.6 Flash, trailed Muse Spark 1.1 and GLM-5.2 by one point, and trailed Kimi K3 by seven. The comparison describes Artificial Analysis’s benchmark mix and settings, not every workload, modality, or deployment requirement.

The pricing makes V4-Flash worth workload-specific testing, especially for high-volume text inference. It is not enough for enterprise approval by itself. Teams still need representative evaluations, accepted-output thresholds, security and data reviews, latency tests, service-level checks, provider assessment, and a measured fallback plan.

Share
PK
Written by
Priyanshi Kharwade
Priyanshi Kharwade — B2B News & Content | Ivris Tech
Content writer covering B2B news and market trends. Communication student with a background in digital marketing and editorial writing. Tracks the developments that matter for B2B operators.

Get B2B marketing insights weekly

Strategies, frameworks, and tools — no fluff. Join operators who read Ivris Tech.

No spam. Unsubscribe anytime.
Link copied!