Close Menu
WithO2WithO2

    Subscribe to Updates

    Get the latest AI News Tools Updates in your Inbox

    What's Hot

    GLM-5.3 Tops Open-Source Coding — and Found 2,436 Security Bugs

    August 16, 2026

    Wix Symphony Gives Small Businesses a Team of AI Agents

    August 16, 2026

    DeepSeek Harness v0.1: The Open-Source AI Agent Runtime Challenging Claude Code

    August 15, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    WithO2WithO2
    • AI
    • Blog
    • Business Software
    • Trending News
    • Stories
    WithO2WithO2
    Home » Trending News
    Trending News

    GLM-5.3 Tops Open-Source Coding — and Found 2,436 Security Bugs

    By Amitabh SarkarAugust 16, 20264 Mins Read0
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Published: August 14, 2026

    Z.ai released GLM-5.3 on August 14, 2026, claiming the strongest open-weights coding model it has measured and disclosing that the model developed multi-step exploit-chain reasoning its training never planned for. Z.ai is withholding the open weights for approximately two weeks while it completes a safety evaluation — the first delayed weight release in the GLM series.

    GLM-5.3 is post-trained on the same 743B-parameter base as GLM-5.2, so every reported gain comes from post-training rather than a larger model. Z.ai reports a 50% improvement in coding capability over GLM-5.2 in internal evaluations, first place among open-source models on Terminal Bench 3.0 and Agents’ Last Exam, and a score of 84.5% on CyberGym, Z.ai’s own cybersecurity evaluation suite. Working alongside security teams, GLM-5.3 identified 2,436 vulnerabilities across 269 projects. The API and the GLM Coding Plan are live now with low, high, and max reasoning levels.

    Table of Contents

    Toggle
    • What a 50% Coding Jump From Post-Training Alone Means
    • The Exploit-Chain Reasoning Z.ai Says It Did Not Plan
    • Why Two Weeks of Safety Hardening Is the Real Signal
    • For Context: Open-Source Coding Models on WithO2

    What a 50% Coding Jump From Post-Training Alone Means

    GLM-5.3 extracts frontier-class coding performance from an unchanged 743B base, which shifts the cost equation for open-weights adopters. Teams that already run GLM-5.2 infrastructure face a capability upgrade rather than a hardware upgrade, because the parameter count and serving footprint are identical.

    Z.ai describes GLM-5.3’s coding and agentic capabilities as “approaching” Claude Fable 5 — near-frontier, not above it on all benchmarks. The two benchmarks Z.ai leads among open-source models, Terminal Bench 3.0 and Agents’ Last Exam, both measure multi-step agentic execution rather than single-shot code completion. That is the workload profile that matters for autonomous development tasks, such as repository-wide refactors, CI failure triage, and dependency migrations. For a benchmark-based comparison of the models businesses deploy for that class of work, see our guide to the best AI agents for business tasks.

    Z.ai has not published GLM-5.3 pricing. The $1.40 and $4.40 per-million-token rates on the GLM-5.2 pricing table are not confirmed for GLM-5.3, and buyers evaluating total cost should treat the rate as unpublished until Z.ai updates it.

    The Exploit-Chain Reasoning Z.ai Says It Did Not Plan

    GLM-5.3’s cybersecurity capability grew faster and further than Z.ai’s training intended, according to the company’s disclosure. The model began chaining individual vulnerability findings into complete, coherent attack sequences — a behavior Z.ai did not target during post-training.

    “GLM-5.3 began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains.” — Z.ai safety team, via TechTimes

    The 84.5% CyberGym score places GLM-5.3 slightly above Mythos 5 and GPT-5.6 Sol on that evaluation. CyberGym is Z.ai’s internal suite, not an industry-standard benchmark, so the comparison is a vendor claim rather than a third-party result. The 2,436 vulnerabilities across 269 projects were found in collaboration with security teams, which makes the finding the more auditable of the two claims.

    Why Two Weeks of Safety Hardening Is the Real Signal

    Every prior GLM release shipped open weights on day one. Holding GLM-5.3’s weights for roughly two weeks of safety evaluation and hardening marks the first time a Chinese open-weights lab has delayed a flagship release explicitly for safety review — a posture enterprise procurement teams already apply to Anthropic and Google DeepMind. Security-sensitive buyers evaluating open-weights models now have a documented safety process to assess, not only a benchmark table.

    For Context: Open-Source Coding Models on WithO2

    • GLM-5.2 review — the 743B base GLM-5.3 is post-trained on, its 1M-context window, and its published pricing.
    • Meta Muse Code — the coding agent that launched at 21x cheaper to try, and GLM-5.3’s closest competitor on price-led adoption.
    • DeepSeek V4-Flash agent benchmarks — the other Chinese open-weights lab competing on agentic coding scores this quarter.
    • AI coding costs — why per-token rates decide open-weights adoption more than benchmark rank.
    Our Take
    The benchmark claims will be re-litigated the moment the weights land; the safety delay is the part that changes behaviour. A Chinese open-weights lab holding a flagship release for hardening is a procurement signal aimed squarely at enterprise buyers who have used “no published safety process” as the reason to stay on closed US APIs. That argument is now weaker. For teams with security workloads, GLM-5.3’s exploit-chain capability is worth evaluating on your own code the week the weights ship — with the caveat that the same capability is what delayed them. Do not commit to a migration on unpublished pricing.
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Amitabh Sarkar
    • Website

    I am a software engineer, I have a passion for working with cutting-edge technologies and staying up-to-date with the latest developments in the field. In my articles, I share my knowledge and insights on a range of topics, including business software, how to set up tools, and the latest trends in the tech industry.

    Related Posts

    Wix Symphony Gives Small Businesses a Team of AI Agents

    August 16, 2026

    DeepSeek Harness v0.1: The Open-Source AI Agent Runtime Challenging Claude Code

    August 15, 2026

    Gemini 3.7 Flash Beats Claude Sonnet 5 on Coding — at Half the Price

    August 15, 2026

    Comments are closed.

    Don't Miss

    Wix Symphony Gives Small Businesses a Team of AI Agents

    By Amitabh SarkarAugust 16, 2026

    Published: August 11, 2026 Wix launched Symphony on August 11, 2026, a standalone multi-agent AI…

    DeepSeek Harness v0.1: The Open-Source AI Agent Runtime Challenging Claude Code

    August 15, 2026

    Gemini 3.7 Flash Beats Claude Sonnet 5 on Coding — at Half the Price

    August 15, 2026

    Lovable Raises $400M: The AI App Builder Fortune 500 Runs On

    August 14, 2026

    Subscribe to Updates

    Get the latest creative news from SmartMag about art & design.

    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    Our Picks

    12 Best Project Management Software Tools in 2026

    August 1, 2026

    9 Best Ecommerce Platforms Compared (2026)

    July 30, 2026

    Rippling vs Gusto vs BambooHR: Full HRMS Comparison 2026

    July 15, 2026

    Best Ecommerce Platform 2026: Top 10 Options Compared

    July 5, 2026
    Editors Picks

    Wix Symphony Gives Small Businesses a Team of AI Agents

    August 16, 2026

    DeepSeek Harness v0.1: The Open-Source AI Agent Runtime Challenging Claude Code

    August 15, 2026

    Gemini 3.7 Flash Beats Claude Sonnet 5 on Coding — at Half the Price

    August 15, 2026

    Lovable Raises $400M: The AI App Builder Fortune 500 Runs On

    August 14, 2026
    About Us
    About Us

    Your Source for Innovation: Discover in-depth guides, solutions, and tools tailored to modern business challenges.

    Links
    • Blog
    • Privacy Policy
    • Contact WithO2.com
    • Terms and Conditions
    Facebook X (Twitter) Instagram Pinterest
    • About
    • Editorial Policy
    • Contact
    • Privacy Policy
    • Terms
    © 2026 WITHO2.COM

    Type above and press Enter to search. Press Esc to cancel.