Hermes 4 405B
Open weightNous Research · Released August 2025 · Knowledge cutoff August 2024
Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with...
Facts and pricing
- Provider
- Nous Research
- Context window
- 131K tokens
- Input price
- 11.0 credits / 1M tokens ($1.00 raw)
- Output price
- 33.0 credits / 1M tokens ($3.00 raw)
- Vision (image input)
- No
- Tool calling
- No
- Extended reasoning
- No
Credits are what Zeplik bills: 1 credit = $0.10, computed from the raw provider rate with a 1.10x margin. Raw prices shown per 1M tokens.
Try Hermes 4 405B now
What Hermes 4 405B is best for
- Everyday chat and drafting on an open-weight model with transparent lineage
Nous Research family
| Model | Context | Input cr/M | Output cr/M | Released |
|---|---|---|---|---|
| Hermes 4 405B | 131K | 11.0 | 33.0 | August 2025 |
| Hermes 3 70B Instruct | 131K | 7.7 | 7.7 | August 2024 |
| Hermes 3 405B Instruct | 131K | 11.0 | 11.0 | August 2024 |
Frequently asked questions
- What is Hermes 4 405B?
- Hermes 4 405B is an AI model by Nous Research. Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with...
- How much does Hermes 4 405B cost on Zeplik?
- Input tokens cost 11.0 credits per million and output tokens 33.0 credits per million (1 credit = $0.10; the raw provider rates are $1.00 and $3.00 per million). New accounts start with free credits.
- How long can a conversation with Hermes 4 405B be?
- Hermes 4 405B has a 131K-token context window (131,072 tokens), which covers the conversation plus any documents you attach.
- Does Hermes 4 405B support images and tools?
- Hermes 4 405B is text-only and does not support tool calling.
Related models
Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for...
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
More you can do with Hermes 4 405B
- 465 AI skillsReady-to-run expert methods for writing, data, code and more, each running on Hermes 4 405B or any model you pick.
- 997 integrationsConnect Gmail, Slack, GitHub, Notion and hundreds more so the assistant can work in your apps.
- Simple pricingPlus, Pro and Max plans with a monthly credit allowance that works across every model.