Original Reddit post

Post-training alone did the heavy lifting on Z.ai’s latest release, and that is the part worth pausing on. In the z.ai launch post, the lab said GLM-5.3 runs on the same mixture-of-experts base as GLM-5.2 and every reported gain came from extended post-training rather than a fresh pretrain. The scoreboard the company is putting out is aggressive: Terminal-Bench 3.0 climbs from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and CyberGym reaches 84.5%, edging Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. The security numbers are what make this launch different from the routine coding-benchmark press release. Z.ai says the model surfaced 2,436 vulnerabilities across 269 open-source projects during evaluation, with 1,097 rated critical or high severity, and reports finding critical bugs in Linux, WebKit, and FreeBSD. The lab also says the model began reasoning across multiple stages of exploitation and forming coherent plans for complete exploitation chains, a capability it did not set out to train for. That admission is why weights are being held back roughly two weeks for safety evaluation and hardening, according to reporting from SiliconANGLE and The Agent Report. For a lab whose open-weight releases are much of the reason its models get attention, that is not a small choice. Some caution is warranted on the specifics. Every score above comes from Z.ai’s own report on its own benchmark mix, so independent reruns have not landed yet, and coverage notes GLM-5.3 still trails Fable 5 and GPT-5.6 Sol badly on ExploitBench and ExploitGym. The launch post does not describe what criteria decide whether the weights actually ship in two weeks or who signs off, and nothing in the reporting pins down whether the 2,436 disclosed bugs were coordinated with the affected maintainers first.

submitted by /u/Justgototheeffinmoon

Originally posted by u/Justgototheeffinmoon on r/ArtificialInteligence