Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks
Z.ai launched GLM-5.3, keeping the 743B base model unchanged but achieving significant improvements through scaled post-training. The model shows massive leaps in coding and cybersecurity benchmarks.

- GLM-5.3 improves performance via post-training only, keeping the base model static.
- Terminal-Bench scores surged from 4.6 to 28.3, indicating better complex coding skills.
- Cybersecurity benchmarks showed significant gains, with ExploitBench doubling to 54.4%.
- Open weights for the model will be released in two weeks.
Z.ai has introduced GLM-5.3, an upgrade that retains the original 743 billion parameter base model from GLM-5.2. Instead of retraining the foundation, the team focused entirely on scaling post-training processes. This involved expanding the variety and duration of task environments to refine the model's capabilities without altering the core weights.
The results are particularly strong in coding and long-horizon tasks. On Terminal-Bench 3.0, the score jumped dramatically from 4.6 to 28.3, while DeepSWE v1.1 improved from 46.2 to 66.9. These figures suggest a substantial leap in the model's ability to handle complex, multi-step programming challenges.
Cybersecurity performance also exceeded expectations. CyberGym reached 84.5%, and ExploitBench scores more than doubled to 54.4%. Z.ai confirmed that the model weights will be made available to the public in approximately two weeks.
Provides access to a high-performing coding model with open weights soon.
Signals potential for more efficient AI updates and better automated security tools.
Demonstrates cost-efficient model improvement strategies without full retraining.
Shows rapid progress in AI capabilities for coding and security tasks.
- Post-training
- The phase after a model is built where it is fine-tuned for specific tasks or safety.
- Long-horizon tasks
- Complex problems that require many steps or a long duration to solve.
- Weights
- The internal numerical parameters of a neural network that determine its behavior.
LLMAlibaba's Qwen team releases Qwen 3.8 models with open weights under the Apache 2.0 license
Meta’s ‘open’ AI, and a $250M deal gone very wrong
LLMGemini 3.7 Flash lands with coding gains and undercuts its three-week-old predecessor's price by 50%
LLMGoogle AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens
LLMGoogle announces Gemini 3.7 Flash just three weeks after previous release
BusinessOne in five US workers now delegates tasks to AI instead of colleagues, survey finds
A new survey reveals that one in five employed Americans delegate at least one work task to AI instead of coworkers, often accepting the output without edits.
AI is changing medicine – but who will pay for the next mistake? - The Jerusalem Post
A Jerusalem Post article examines who bears financial responsibility when AI-driven medical tools cause errors.
Clinical Applications, Opportunities, and Implementation Challenges of AI in Emergency Medicine: A Narrative Review - Cureus
A narrative review in Cureus explores the clinical applications, opportunities, and implementation challenges of AI in emergency medicine.
Home province of DeepSeek, Moonshot founders seeks to retain future AI talent - South China Morning Post
The home province of DeepSeek and Moonshot founders is taking steps to retain future AI talent, according to a recent report.
SecurityI shipped an MCP server that reported success without signing anything
A developer built an MCP server that simulated successful token trades on Solana without requiring cryptographic signatures, raising security concerns.
Will AI data centers soak up the last of the valley’s water? - Fresno Bee
A Fresno Bee investigation warns that AI data centers in California's Central Valley could exacerbate water scarcity amid ongoing drought conditions.