Briefing

DeepSeek Announces v4 Models, Pricing Adjustments, and New Features

ai-dev
by nateb2022 · DeepSeek

Check the updated DeepSeek pricing and feature set, and adjust your token cost calculations accordingly.

What to do now

Patch your cost estimation scripts to use the new DeepSeek pricing and update your model selection logic to prefer deepseek-v4-flash for non‑thinking workloads.

Summary

DeepSeek released two new models, deepseek-v4-flash and deepseek-v4-pro, with a 1M token context length and a 384K token maximum output. They support both thinking and non‑thinking modes, JSON output, tool calls, chat prefix completion (beta), and FIM completion (beta). The new models replace the legacy deepseek-chat and deepseek-reasoner names, which will be deprecated.

Pricing changes effective 2026/04/26 see the input cache hit price dropped to 1/10 of the launch price, and deepseek-v4-pro is offered at a 75% discount until 2026/05/31 15:59 UTC. Input cache miss and output token prices remain unchanged but are listed per 1M tokens. The announcement notes that fees are deducted from topped‑up or granted balances, preferring granted when both exist. The new models also introduce a thinking mode toggle, allowing developers to choose between non‑thinking (default) and thinking (advanced) behavior, and the structured output capabilities support integration with external APIs.

Key changes

  • Introduction of deepseek-v4-flash and deepseek-v4-pro models with 1M context length and 384K max output
  • Support for both thinking and non‑thinking modes, replacing legacy deepseek-chat and deepseek-reasoner names
  • New structured output capabilities: JSON output, tool calls, chat prefix completion (beta), and FIM completion (beta)
  • Pricing change effective 2026/04/26: input cache hit price reduced to 1/10 of launch price
  • Deepseek-v4-pro offered at 75% discount until 2026/05/31 15:59 UTC
  • Input cache miss and output token prices remain unchanged but are listed per 1M tokens
  • Fees deducted from topped‑up or granted balances, preferring granted when both exist

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting