GLM-5.2
Flagship open reasoning model for agentic software engineering with 256K context window
About model
GLM-5.2 is Z.ai's newest flagship model for agentic coding and long-horizon software engineering, built on the GLM-5 Mixture-of-Experts backbone — 744B total parameters with 40B active per token. The release also adds two thinking-effort levels, letting developers trade compute depth against throughput depending on task complexity. GLM-5.2 ships under the MIT License with open weights. According to Z.ai, the model delivers powerful coding capabilities and continued strengths in long-horizon agentic tasks.
256K
Usable 256K-token context window with a 131,072-token output cap
744B
Mixture-of-Experts backbone optimized for coding agents
Configurable
Two thinking-effort levels exposed directly at the API level
- Long-Context Agentic Coding: A usable 256K-token context window — 5× larger than GLM-5.1's ~200K limit — with a 131,072-token output cap. According to Z.ai, the model tracks cross-file dependencies across full repositories in a single pass, enabling whole-repo refactors without repeated context retrieval.
- Configurable Thinking Effort: Two thinking-effort levels exposed at the API level let developers dial between lighter inference for fast tasks and deeper reasoning for complex long-horizon sessions — a tradeoff between throughput and reasoning depth.
- Drop-in Agent Compatibility: Day-one support for eight coding agents including Claude Code, Cline, Roo Code, Goose, and OpenCode via an OpenAI-compatible and Anthropic Messages API — a config-level swap with no framework changes required.
Model | FrontierMath Tier 4 | GPQA Diamond | HLE | SciCode | GDPval-AA | Terminal-Bench 2.1 | Agent Arena | FrontierCode | DeepSWE |
|---|---|---|---|---|---|---|---|---|---|
GLM-5.2 | 29.3% | 89.5% | 40% | 50% | 50% | 78% | +9.1pp | 24.5% | 44% | Related open-source models | Competitor closed-source models |
87.8% | 92.6% | 53% | 60% | 62% | 85% | 53.5% | 70% | ||
73.2% | 93.2% | 53% | 56% | 68% | 89% | +16.4pp | 53.4% | 74% | |
82.9% | 94.1% | 47% | 56% | 61% | 88% | 47.5% | 73% | ||
24.4% | 93.1% | 40% | 54% | 51% | 82% | +4.2pp | 42.4% | 54% | |
61.0% | 91.1% | 37% | 53% | 54% | 81% | 39.8% | 67% |
API usage
Endpoint:
Model card
Architecture Overview:
• 744B total parameter MoE architecture with 40B parameters activated per token
• 256K context window with 131,072 max output tokens
• Supports configurable thinking modes (two effort levels) for step-by-step reasoning
• Open-weights distribution under the MIT License
Training Methodology:
• Pre-trained and instruction-tuned entirely on Huawei Ascend 910B hardware under the MindSpore framework
• Post-training optimizations targeting repo-scale dependencies and extended plan-then-execute loops
Performance Characteristics:
• Flagship reasoning performance built specifically for long-horizon orchestration
• Native framework compatibility with multiple coding agent workflows
• No official benchmarks were published at launch; performance metrics will be updated as they are verified
Prompting
Together AI API Access:
• Access GLM-5.2 via Together AI APIs using the endpoint zai-org/GLM-5.2
• Authenticate using your Together AI API key in request headers
• Supports configurable thinking effort parameters
• Available on Together AI serverless infrastructure
Applications & use cases
Repository-Scale Engineering:
• Complete multi-file refactoring, code migration, and complex system updates in a single pass
• Navigation of large codebases without needing vector-search retrieval chunking
Autonomous Technical Workflows:
• Extended plan-then-execute cycles utilizing up to 131K output tokens for deep tracing
• Multi-step tool use, system debugging, and algorithmic optimization workflows
Drop-in Agent Integrations:
• Full compatibility with Claude Code, Cline, Roo Code, Goose, OpenCode, and custom frameworks
• Switch configurations smoothly via OpenAI-compatible and Anthropic Messages API formats
- TypeReasoningChatCode
- Main use casesReasoning
- FeaturesFunction CallingJSON Mode
- DeploymentServerlessProvisioned Throughput
- Endpoint
- Parameters753B
- Activated parameters40B
- Context length256K
- Input price
$1.40 / 1M tokens
$0.26 (cached)/1M
- Output price
$4.40 / 1M tokens
- Input modalitiesText
- Output modalitiesText
- ReleasedJune 16, 2026
- Quantization levelFP4
- CategoryChat
