Z.AI Unveils GLM-4.6: 355B Parameter MoE Model with Enhanced Coding Capabilities
Zhipu AI releases GLM-4.6 with 355B parameter MoE architecture, 200K token context window, and superior coding performance in 74 real-world tests.
Zhipu AI (also know as z.AI) has unveiled GLM-4.6, a groundbreaking large language model featuring a 355B parameter Mixture of Experts (MoE) architecture that represents a significant leap forward in AI capabilities. The model expands the context window to 200K tokens, doubling the capacity of its predecessor GLM-4.5’s 128K token limit, enabling more complex and extended interactions.
In comprehensive performance evaluations, GLM-4.6 demonstrates remarkable coding prowess, outperforming competitors in 74 real-world coding tests conducted within the Claude Code environment. The model achieves over 30% greater token efficiency compared to other leading models, consuming approximately 15% fewer tokens than GLM-4.5 while delivering superior performance.
Benchmark results place GLM-4.6 on par with Claude Sonnet 4 and 4.5 across multiple authoritative evaluations, though it shows competitive advantages over Claude Sonnet 4 while lagging slightly behind Claude Sonnet 4.5 in coding capabilities. The model has been evaluated across eight public benchmarks including AIME 25, GPQA, LCB v6, HLE, and SWE-Bench Verified, demonstrating strong multi-disciplinary reasoning and coding performance.
Zhipu AI CEO Zhang Peng has expressed optimism about the AGI timeline, anticipating the emergence of Artificial General Intelligence around 2025. This vision aligns with the rapid advancements demonstrated by GLM-4.6, which represents a significant step toward more sophisticated AI systems.
The model’s architecture enables efficient computation by leveraging only a subset of its 355B parameters for each task, making it both powerful and computationally efficient. With a maximum output token limit of 128K tokens, GLM-4.6 provides ample capacity for detailed and elaborate responses across various applications.
GLM-4.6’s enhanced capabilities extend beyond coding to include stronger performance in tool use and search-based agents, improved cross-lingual task performance, and better alignment with human preferences in style, readability, and role-playing scenarios. The model’s release in 2024 marks a significant milestone in the evolution of large language models, setting new standards for efficiency and performance in real-world applications.
The new model can be tried locally, via API or via the https://chat.z.ai/ .