Skip to main content

DeepSeek models

DeepSeek publishes open-weight models that compete with far more expensive closed models, particularly on reasoning and code. Zeplik carries the V4 generation (Pro and Flash) and the V3.x line behind it.

All 16 DeepSeek models

DeepSeek V4 Flash Vision ExpFast and light

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

1.0M ctx7.3 cr/M outAugust 2026
DeepSeek V4 Pro 0813Open weight

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

1.0M ctx19.1 cr/M outAugust 2026
DeepSeek V4 Pro 0813 (batch)Open weight

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

1.0M ctx43.6 cr/M outAugust 2026
DeepSeek V4 Flash 0731Fast and light

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

1.0M ctx2.0 cr/M outJuly 2026
DeepSeek V4 Flash 0731 (batch)Fast and light

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

1.0M ctx3.1 cr/M outJuly 2026
DeepSeek V4 Pro 0423Open weight

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

1.0M ctx18.2 cr/M outApril 2026
DeepSeek V4 Flash 0423Fast and light

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

1.0M ctx1.9 cr/M outApril 2026
DeepSeek V3.2Open weight

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

164K ctx4.4 cr/M outDecember 2025
DeepSeek V3.2 ExpOpen weight

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

164K ctx4.5 cr/M outSeptember 2025
DeepSeek V3.1 TerminusOpen weight

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...

131K ctx11.0 cr/M outSeptember 2025
DeepSeek V3.1Open weight

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

161K ctx18.2 cr/M outAugust 2025
R1 0528Open weight

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

164K ctx23.7 cr/M outMay 2025
DeepSeek V3 0324Open weight

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...

164K ctx11.0 cr/M outMarch 2025
R1 Distill Llama 70BOpen weight

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

8K ctx8.8 cr/M outJanuary 2025
R1Open weight

DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....

64K ctx27.5 cr/M outJanuary 2025
DeepSeek V3Open weight

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations...

164K ctx9.8 cr/M outDecember 2024

Pricing at a glance

ModelContextInput cr/MOutput cr/MReleased
DeepSeek V4 Flash Vision Exp1.0M2.47.3August 2026
DeepSeek V4 Pro 08131.0M6.419.1August 2026
DeepSeek V4 Pro 0813 (batch)1.0M14.543.6August 2026
DeepSeek V4 Flash 07311.0M0.722.0July 2026
DeepSeek V4 Flash 0731 (batch)1.0M1.53.1July 2026
DeepSeek V4 Pro 04231.0M9.118.2April 2026
DeepSeek V4 Flash 04231.0M0.931.9April 2026
DeepSeek V3.2164K3.04.4December 2025
DeepSeek V3.2 Exp164K3.04.5September 2025
DeepSeek V3.1 Terminus131K3.011.0September 2025
DeepSeek V3.1161K6.018.2August 2025
R1 0528164K5.523.7May 2025
DeepSeek V3 0324164K2.811.0March 2025
R1 Distill Llama 70B8K8.88.8January 2025
R164K7.727.5January 2025
DeepSeek V3164K3.59.8December 2024

Other providers

Explore Zeplik

DeepSeek Models - Pricing, Context and Capabilities | Zeplik Chat