Skip to main content

Z.ai models

Z.ai (formerly Zhipu) ships the GLM series, open-weight models with strong coding and agent performance. Zeplik carries the GLM 5.x generation including Turbo speed tiers and the vision-capable 5V.

All 16 Z.ai models

GLM 5.3 Flash (batch)Fast and light

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

1.0M ctx5.5 cr/M outAugust 2026
GLM 5.3 FlashFast and light

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

1.0M ctx2.8 cr/M outAugust 2026
GLM 5.3Open weight

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

1.0M ctx48.4 cr/M outAugust 2026
GLM 5.2Open weight

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

1.0M ctx33.4 cr/M outJune 2026
GLM 5.2 (free)Open weight

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

256K ctx0 cr/M outJune 2026
GLM 5.1Open weight

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

200K ctx33.4 cr/M outApril 2026
GLM 5V TurboFast and light

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

203K ctx44.0 cr/M outApril 2026
GLM 5 TurboFast and light

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...

203K ctx44.0 cr/M outMarch 2026
GLM 5Open weight

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...

198K ctx21.1 cr/M outFebruary 2026
GLM 4.7 FlashFast and light

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

203K ctx4.4 cr/M outJanuary 2026
GLM 4.7Open weight

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

203K ctx19.3 cr/M outDecember 2025
GLM 4.6VOpen weight

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

131K ctx9.9 cr/M outDecember 2025
GLM 4.6Open weight

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

205K ctx24.2 cr/M outSeptember 2025
GLM 4.5VOpen weight

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...

66K ctx19.8 cr/M outAugust 2025
GLM 4.5Open weight

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

131K ctx24.2 cr/M outJuly 2025
GLM 4.5 AirFast and light

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

131K ctx9.3 cr/M outJuly 2025

Pricing at a glance

ModelContextInput cr/MOutput cr/MReleased
GLM 5.3 Flash (batch)1.0M1.75.5August 2026
GLM 5.3 Flash1.0M0.832.8August 2026
GLM 5.31.0M15.448.4August 2026
GLM 5.21.0M10.633.4June 2026
GLM 5.2 (free)256K00June 2026
GLM 5.1200K10.633.4April 2026
GLM 5V Turbo203K13.244.0April 2026
GLM 5 Turbo203K13.244.0March 2026
GLM 5198K6.621.1February 2026
GLM 4.7 Flash203K0.664.4January 2026
GLM 4.7203K4.419.3December 2025
GLM 4.6V131K3.39.9December 2025
GLM 4.6205K6.024.2September 2025
GLM 4.5V66K6.619.8August 2025
GLM 4.5131K6.624.2July 2025
GLM 4.5 Air131K1.49.3July 2025

Other providers

Explore Zeplik

Z.ai Models - Pricing, Context and Capabilities | Zeplik Chat