Skip to main content

R1 Distill Llama 70B

Open weight

DeepSeek · Released January 2025 · Knowledge cutoff July 2024

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

It is part of DeepSeek's open-weight line, known for delivering near-frontier reasoning and coding quality at a fraction of typical frontier pricing.

Facts and pricing

Provider
DeepSeek
Context window
8K tokens
Input price
8.8 credits / 1M tokens ($0.80 raw)
Output price
8.8 credits / 1M tokens ($0.80 raw)
Vision (image input)
No
Tool calling
No
Extended reasoning
No

Credits are what Zeplik bills: 1 credit = $0.10, computed from the raw provider rate with a 1.10x margin. Raw prices shown per 1M tokens.

Try R1 Distill Llama 70B now

Ask R1 Distill Llama 70B anything. Your prompt opens in the Zeplik app with this model selected.

R1 Distill Llama 70B

Or open Zeplik with R1 Distill Llama 70B already selected

What R1 Distill Llama 70B is best for

Example prompts

Prompts that suit a open weight model like R1 Distill Llama 70B:

DeepSeek family

ModelContextInput cr/MOutput cr/MReleased
DeepSeek V4 Flash Vision Exp1.0M2.47.3August 2026
DeepSeek V4 Pro 08131.0M6.419.1August 2026
DeepSeek V4 Pro 0813 (batch)1.0M14.543.6August 2026
DeepSeek V4 Flash 07311.0M0.722.0July 2026
DeepSeek V4 Flash 0731 (batch)1.0M1.53.1July 2026
DeepSeek V4 Pro 04231.0M9.118.2April 2026
DeepSeek V4 Flash 04231.0M0.931.9April 2026
DeepSeek V3.2164K3.04.4December 2025
DeepSeek V3.2 Exp164K3.04.5September 2025
DeepSeek V3.1 Terminus131K3.011.0September 2025

All 16 DeepSeek models

Frequently asked questions

What is R1 Distill Llama 70B?
R1 Distill Llama 70B is an AI model by DeepSeek. DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...
How much does R1 Distill Llama 70B cost on Zeplik?
Input tokens cost 8.8 credits per million and output tokens 8.8 credits per million (1 credit = $0.10; the raw provider rates are $0.80 and $0.80 per million). New accounts start with free credits.
How long can a conversation with R1 Distill Llama 70B be?
R1 Distill Llama 70B has a 8K-token context window (8,192 tokens), which covers the conversation plus any documents you attach.
Does R1 Distill Llama 70B support images and tools?
R1 Distill Llama 70B is text-only and does not support tool calling.

Related models

DeepSeek V4 Flash Vision ExpFast and light

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

1.0M ctx7.3 cr/M outAugust 2026
DeepSeek V4 Pro 0813Open weight

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

1.0M ctx19.1 cr/M outAugust 2026
DeepSeek V4 Pro 0813 (batch)Open weight

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

1.0M ctx43.6 cr/M outAugust 2026
DeepSeek V4 Flash 0731Fast and light

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

1.0M ctx2.0 cr/M outJuly 2026
DeepSeek V4 Flash 0731 (batch)Fast and light

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

1.0M ctx3.1 cr/M outJuly 2026
DeepSeek V4 Pro 0423Open weight

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

1.0M ctx18.2 cr/M outApril 2026

More you can do with R1 Distill Llama 70B

R1 Distill Llama 70B - DeepSeek AI Model | Zeplik Chat