AI models on Zeplik
289 models from 50 labs, all in one chat. Pricing and capabilities below come straight from the live registry Zeplik routes and bills against, synced continuously. Pick a model to see its full profile, or start typing on any model page to try it.
Featured
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work.
GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier.
Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5.
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work.
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series.
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier.
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6.
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window.
Media generation
Beyond text, Zeplik generates images, video, music and speech — and transcribes audio — right in the chat. You never pick a media model; just ask, and the product routes to the right one. Costs below are the exact credit basis the ledger charges.
Video
Fast (default)Kling 2.5
~0.84 credit / second
a 5-second clip ≈ 4.2 credits (up to ~12.6 at the 15s max)
Video
Premium (cinematic/HD)Kling 2.6
~1.68 credit / second
a 10-second clip ≈ 16.8 credits (up to ~25.2 at the 15s max) — asks you to confirm the spend first
Confirms the spend before generating
Text-to-speech
xAI TTS
~0.18 credit / 1,000 characters
about 0.22 credits for a reply-length passage
Music
CassetteAI
~0.24 credit / minute (rounded up)
a 30-second track ≈ 0.24 credits
3D model (from text)
Rodin
~4.8 credits / mesh
describe an object and get a rotatable GLB
3D model (from an image)
Hunyuan3D 2.1
~3.6 credits / mesh
turn a photo in the conversation into a rotatable GLB
Image (generate + edit)
Nano Banana 2
~0.96 credit / image
generate a new image or edit one in the conversation
Upscale
Clarity Upscaler
~0.36 credit / image
increase the resolution of an image in the conversation
Background removal
BiRefNet
~0.01 credit / image (a fraction of a cent)
remove the background, leaving a transparent PNG
Vectorize
Recraft
~0.12 credit / image
convert an image into a scalable SVG vector
Vector/SVG from text
Recraft
~0.96 credit / image
generate a brand-new SVG (logo, icon) from a description
Transcription
Whisper / Wizper
Free
upload an audio file and its speech is transcribed to text
OpenAI
OpenAI ships the GPT line, the most widely used family of AI models in the world. Zeplik leads with GPT-6 Astra and its deeper Astra Pro serving, and now carries the rest of the GPT-6 generation beneath it: GPT-6 Sol for demanding professional work and GPT-6 Luna for high-volume, latency-sensitive work, each with a Pro serving and each on the same 1M-token window as Astra. Behind them sit the GPT-5.6 tiers (Sol, Luna and Terra), GPT-5 and GPT-4 era releases, mini and nano tiers for fast inexpensive work, the Codex models for software engineering, GPT Audio for speech, and the GPT Image models for generation and editing.
GPT-6.1 Sol Pro is the same underlying model as [GPT-6.1 Sol](https://openrouter.ai/openai/gpt-6.1-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks.
GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series.
GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol.
GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol.
Qwen
Alibaba's Qwen team ships one of the broadest open catalogs in AI: Max and Plus capability tiers, Flash speed tiers, dedicated coder and reasoning variants, vision-language models, and dozens of open-weight sizes. Zeplik carries the Qwen3.x generation led by Qwen3.8 Max and its higher-throughput Max Prime serving, with Qwen3.8 Flash and the multimodal Omni Flash on the speed tier and the long tail of open-weight sizes behind them.
Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point.
Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding.
Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team.
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.
Qwen3.8 27B is an open-weight dense vision-language model from Qwen.
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total.
Google DeepMind's Gemini models pair strong multimodal understanding with some of the largest context windows available, and the open-weight Gemma line brings the same research to smaller, cheaper models. Zeplik carries the Gemini 3.x generation led by Gemini 3.8 Flash, the Flash-Lite speed tiers, the Nano Banana image models, Lyria for music, and Gemma 4.
Nano Banana 2.1 (Gemini Nano Banana 2.1) is Google's image generation and editing model on the Flash tier, succeeding Nano Banana 2 and Nano Banana Pro.
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning.
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning.
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development.
Anthropic
Anthropic builds the Claude family, known for careful reasoning, strong writing and reliable tool use. Zeplik leads with Claude Opus 5.5, Anthropic's flagship for demanding reasoning, coding and long-horizon agentic work, which succeeds Claude Opus 5 at a lower price on both input and output. Alongside it Zeplik carries Claude Fable 5.1 and the rest of the current generation — Claude Opus 5, Claude Sonnet 5 and Claude Fable 5 — with the Opus 4.x line and earlier Sonnet and Haiku releases behind them.
Claude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use.
Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade.
Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5.
Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5.
Mistral
Mistral AI, the leading European lab, ships efficient models across every size. Zeplik carries Mistral Medium 3.5 at the head of the catalog, plus Mistral Large 3 and Mistral Small 4, the Ministral 3 sizes for edge deployment, Devstral 2 and Codestral for code, and Voxtral for audio.
Mistral Large 4 is a frontier multimodal (text and image input) model from Mistral AI built for reasoning, coding, and agentic workloads.
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI.
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI.
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system.
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system.
Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window.
Z.ai
Z.ai (formerly Zhipu) ships the GLM series, open-weight models with strong coding and agent performance. Zeplik carries the GLM 5.x generation led by GLM 5.3 and its high-throughput 5.3 Prime serving, with the 5.3 Flash and FlashX speed tiers, the Turbo servings and the vision-capable 5V behind them.
GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration.
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s.
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks.
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks.
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks.
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks.
DeepSeek
DeepSeek publishes open-weight models that compete with far more expensive closed models, particularly on reasoning and code. Zeplik leads with DeepSeek V4.1 Flash — the first DeepSeek release on the company's Causal Encoder-Decoder architecture, with vision and a 1M-token window — and carries the V4 generation behind it, including Pro, Flash and the vision-capable V4 Flash Vision, along with the V3.x and R1 lines.
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture.
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture.
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total.
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window.
NVIDIA
NVIDIA's Nemotron models are open-weight releases tuned for helpfulness and agentic tasks, built on top of leading open architectures and optimized for NVIDIA hardware. Zeplik carries Nemotron 3.5 Lightning at the head of the line, plus the Nemotron 3 Ultra, Super and Nano sizes and a content-safety classifier.
Switchyard is an open-source model router that switches between multiple models to optimize the cost of requests.
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total.
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total.
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B.
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B.
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE).
SpaceXAI
SpaceXAI — the lab formerly known as xAI — trains the Grok models with an emphasis on reasoning and up-to-date knowledge. Zeplik carries the Grok 4.x generation led by Grok 4.7, which succeeds Grok 4.6 on long-running software engineering and agentic work while costing less on both input and output, with Grok 4.6, Grok 4.5 and Grok 4.3 behind it, the Grok 4.20 Multi-Agent serving, and the coding-focused Grok Build.
Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6.
Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by [Grok 4.7](/x-ai/grok-4.7).
Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows.
Grok 4.3 is a reasoning model from SpaceXAI.
Grok 4.3 is a reasoning model from SpaceXAI.
MoonshotAI
Moonshot AI builds the Kimi models, open-weight releases with standout agentic and coding ability. Zeplik leads with Kimi K3 and carries the K2.x line behind it, including the code-focused K2.7 Code and the deliberate K2 Thinking.
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI.
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI.
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts.
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration.
Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm.
Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning.
MiniMax
MiniMax builds the M-series of chat models, iterating quickly on conversational quality and long-context handling. Zeplik carries MiniMax M3 at the head of the line, with the M2.x releases and the original MiniMax-01 behind it.
MiniMax-M3 is a multimodal foundation model from MiniMax.
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement.
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity.
MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn conversations.
MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development.
MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows.
Meta
Meta's Llama models are the most widely adopted open-weight family in the industry. Zeplik carries the Llama 4 generation (Maverick and Scout), the Llama Guard 4 safety classifier, and the proven Llama 3.x releases.
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification.
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B.
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).
Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis.
Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization.
Tencent
Tencent's Hunyuan models span general chat and dedicated translation. Zeplik carries the Hy4 preview at the head of the line, plus Hy3, the Hy-MT2 translation models in three sizes, and the open-weight Hunyuan A13B.
Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total.
Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent.
Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family.
Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation.
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use.
Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use.
AionLabs
AionLabs ships the Aion models, multi-model roleplaying and storytelling systems in which several specialised models generate collaboratively. Zeplik carries Aion-3.5 and Aion-3.5-Mini at the head of the line, with Aion-3.0, Aion-3.0-Mini, Aion-2.0 and the role-play focused Aion-RP behind them.
Aion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models.
Aion 3.5 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models.
Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models.
Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models.
Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling.
Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of Arena-Hard-Auto, where LLMs evaluate each other’s responses.
Cohere
Cohere focuses on enterprise text work: retrieval-augmented generation, tool use and multilingual business writing. Zeplik leads with Command A+, Cohere's flagship for enterprise agentic workflows, which takes text and image input with strict tool schemas over a 192K window, and carries North Mini Code alongside the rest of the Command family — Command A, Command R+ and Command R.
Command A+ is Cohere's flagship model for enterprise agentic workflows.
North Mini Code is Cohere's first agentic coding model and the debut of its North family.
Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases.
Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024.
command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmented generation (RAG) and tool use.
ByteDance Seed
ByteDance's Seed team ships the Seed series, models trained for strong general reasoning with a dedicated code variant. Zeplik carries Seed 2.1 Turbo and Seed-2.0-Code at the head of the line, with the Seed-2.0 Lite and Mini sizes and the Seed 1.6 releases behind them.
Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows.
Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding.
Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment.
Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding.
Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.
OpenRouter
OpenRouter's own entries are ROUTERS, not models: each one reads the request and forwards it to whichever underlying model fits best, so its price and capabilities depend on where it lands. Zeplik carries the Auto Router and its beta, the Fusion and Pareto Code routers, and the free-models router.
The experimental version of our Auto Router where we test new improvements.
Fusion turns your prompt into a small multi-model deliberation.
The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) coding percentiles.
The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter.
Transform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI models, and Body Builder will construct the appropriate API calls.
The Auto Router automatically selects the best model for your prompt, powered by the wisdom of the market.
inclusionAI
inclusionAI publishes the open-weight Ling models. Zeplik carries the Ling 3.0 Flash line led by Ling 3.0 Flash VL, which adds native visual perception to the base Flash model, along with the domain-tuned Sante (health) and Fin (finance) variants.
Ling 3.1 Flash is a hybrid reasoning mixture-of-experts model from inclusionAI, with 25B active parameters out of 560B total.
Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total.
Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total.
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*.
Xiaomi
Xiaomi's MiMo models are open-weight releases from the company's AI lab. Zeplik carries the MiMo-V2.6 generation — the 1T-parameter MiMo-V2.6-Pro, its roughly 10x faster Pro UltraSpeed serving, and the lighter MiMo-V2.6-Flash — with MiMo-V2.5 and MiMo-V2.5-Pro behind them.
MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro.
MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi.
MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi.
MiMo-V2.5 is a native omnimodal model by Xiaomi.
Perplexity
Perplexity's Sonar models are built for search-grounded answering: they retrieve from the live web and cite sources as part of the response. Zeplik carries Sonar Pro Search at the head of the line, with Sonar Reasoning Pro, Sonar Deep Research and the base Sonar model behind it.
Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis.
Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing-breakdown-for-sonar-reasoning-pro-and-sonar-pro) Sonar Reasoning Pro is a premier reasoning model powered by DeepSeek R1 with Chain of Thought (CoT).
Note: Sonar Pro pricing includes Perplexity search pricing.
Sonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics.
Sonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources.
Thinking Machines
Thinking Machines ships the Inkling models. Zeplik carries both sizes — Inkling and the lighter Inkling Small — each with batch and free servings.
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total.
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total.
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total.
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total.
Poolside
Poolside builds models for software engineering. Zeplik carries the Laguna line in two sizes — Laguna S 2.1 and the smaller Laguna XS 2.1 — each with a free serving.
Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>).
Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>).
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026).
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026).
Amazon
Amazon's Nova models are the in-house family behind AWS Bedrock, tuned for practical business tasks at aggressive price points. Zeplik carries Nova 2 Lite and Nova Premier alongside the original Nova Pro, Lite and Micro tiers.
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text.
Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output.
Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost.
Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks.
Upstage
Upstage builds the Solar models, compact releases that punch above their size on reasoning and document work. Zeplik carries Solar Mini 4 — a 35B mixture-of-experts with 3B active parameters and a 524K window, built for agentic work where response cost matters — alongside Solar Pro 4 and Solar Pro 3.
Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context window.
Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window.
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency.
Sakana
Sakana AI, the Tokyo lab, researches nature-inspired approaches to model design. Its Fugu models are not single monolithic models but a learned multi-agent orchestration system that routes work across specialised models. Zeplik carries Fugu Ultra v2 and the cost-performance Fugu Max, with the original Fugu Ultra behind them.
Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family.
Fugu Max is the cost-performance model in Sakana AI's Fugu family.
Fugu Ultra is the higher-performance model in Sakana AI's Fugu family.
Nous Research
Nous Research fine-tunes open-weight base models into the Hermes series, known for instruction following and a neutral, steerable voice. Zeplik carries Hermes 4 405B and the Hermes 3 releases.
Unbiased
Unbiased builds Pareto, a multimodal composite model for research, coding and agentic workflows. Zeplik carries it with vision, tool calling and a 262K-token window.
Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.
Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.
Perceptron
Perceptron builds models for grounded perception — reading what a scene contains and pointing at it, rather than describing it in prose. Zeplik carries Perceptron Mk1.5, an embodied-reasoning model for physical agents that takes text, images, video and audio and answers with text plus structured annotations — points, boxes, polygons and tracks — with tool calling and adjustable reasoning effort, and Perceptron Mk1 behind it.
Inference.net
Inference.net builds the Schematron models, small models trained for one job: turning HTML into structured JSON against a caller-supplied schema. Zeplik carries both V2 servings — Turbo, tuned for throughput on high-volume extraction, and Small, tuned for extraction quality on complex schemas and long pages.
Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads.
Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages.
Inception
Inception builds diffusion language models (dLLMs), which produce and refine many tokens in parallel rather than one after another — a different generation mechanism from every other lab in this catalog. Zeplik carries Mercury 2.5, the newest of them, with a 260K window and tool calling, alongside Mercury 2.
IBM
IBM's Granite models are open-weight releases aimed at enterprise deployment, published with clear licensing and small enough to run close to the data. Zeplik carries Granite 4.2 8B and the compact Granite 4.0 Micro.
Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning.
Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM.
StepFun
StepFun ships the Step series. Zeplik carries the Flash speed tiers — Step 3.7 Flash and Step 3.5 Flash.
Reka
Reka builds compact multimodal models designed to run efficiently. Zeplik carries Reka Edge and Reka Flash 3.
Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs.
Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat, coding tasks, instruction-following, and function calling.
Morph
Morph builds models for applying code edits fast — taking a proposed change and merging it into a file, rather than writing prose about it. Zeplik carries Morph V3 Large and Morph V3 Fast.
Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations.
Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations.
Microsoft
Microsoft Research publishes the Phi models, small releases trained on carefully curated data that compete with much larger ones on reasoning. Zeplik carries Phi 4 and the WizardLM-2 8x22B mixture-of-experts model.
[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed.
WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary models, and it consistently outperforms all existing state-of-the-art opensource models.
Fireworks
Fireworks Research builds models tuned to make each token count. Zeplik carries Ember-1, a reasoning model built on Kimi K3 that reaches its answer with substantially shorter reasoning traces, over a 1M-token window with vision and tool calling.
PrismML
PrismML works on making capable models small enough to run cheaply. Zeplik carries Ternary Bonsai 2 27B, a reasoning model derived from Qwen3.8-27B and shrunk by ternary compression, which handles code, mathematics, tool calling and image understanding over a 262K window.
Dots Studio
Dots Studio publishes the Dots models. Zeplik carries the Dots3-Note preview with a free serving.
LiquidAI
Liquid AI builds the LFM models on a non-transformer architecture designed for efficiency at small scale. Zeplik carries LFM2.5-2.6B with a free serving.
Meta
Beyond Llama, Meta publishes the Muse line under its own namespace. Zeplik carries Muse Glimmer 30B, with a batch serving.
Meituan
Meituan publishes the open-weight LongCat models. Zeplik carries LongCat 2.0.
Arcee AI
Arcee AI builds open-weight models through model merging and distillation. Zeplik carries Trinity Large Thinking, its reasoning release.
Writer
Writer builds the Palmyra models for enterprise content work — brand-consistent writing, structured documents and business analysis. Zeplik carries Palmyra X5.
Relace
Relace builds models for codebase retrieval — finding the files and symbols a change touches. Zeplik carries Relace Search.
ByteDance
ByteDance publishes UI-TARS, an open-weight model trained specifically to operate graphical user interfaces — reading a screenshot and deciding what to click, type or scroll.
Baidu
Baidu's ERNIE models are a long-running Chinese-language family with strong multilingual and multimodal coverage. Zeplik carries the vision-language ERNIE 4.5 VL, a 424B mixture-of-experts release.
More providers
- TheDrummer (3)
- Sao10K (3)
- Nex AGI (2)
- Apodex (1)
- Typesafe (1)
- Venice (1)
- Anthracite (1)
- Mancer (1)
- Undi95 (1)
- Gryphe (1)
Compare models side by side
Head-to-head pricing, context and capability comparisons for the pairings people actually decide between.