GLM 5.2 on CyberBT-CTF: The strongest open source contender to Anthropic/OpenAI we have tested

Louie.ai researchers found two major surprises in Z.ai’s new GLM 5.2 model:
1. It’s the “Chaotic Good Goblin Paladin” of AI – an open weights Chinese model that runs at half the cost of Anthropic and OpenAI, yet goes toe-to-toe with them to tie on the cheating-resistant CyBT-CTF security agent investigation benchmark.
2. The correct vs wrong results are so statistically similar that we have to ask: Did Z.ai perform the first known successful model distillation attack against frontier model providers? Anthropic previously reported attempts by other Chinese model makers, but Z.ai was not named, and this would be the first publicly measured sign of distillation yielding such frontier-level results.
For those new to Graphistry & Louie.ai, we’ve been running botsbench.com, a continuous evaluation of leading model, harness, and provider combinations on representative agentic cybersecurity investigations. Botsbench is a bit unusual in its extreme care in protecting against benchmark confounds like model contamination, sandbox cheating escapes, and vendor bias. As a plus, it gives comparison points to professional analysts tackling the same challenges. You can find the following numbers on the botsbench GLM 5.2 page. It compares GLM 5.2 used by the OpenCode harness with the Fireworks AI inference provider on the CyBT-CTF and Splunk Botsv3 CTFs blue team agentic investigation benchmarks against our historic database of popular model x harness combos.
The tl;dir of GLM 5.2: With a 28/59 solve rate, it is the top open weight model, and impressively, ties the proprietary ones. While Claude Code / Opus 4.7 does run 19% faster than OpenCode / GLM 5.2, we find Opus to cost 2.2x+ more for the same results. When Cerebras makes GLM 5.2 available, we expect the speed advantage to disappear. Alternative open source models are much worse: The next best open model solve rate is MiniMax 2.5 at 16/59, which GLM 5.2 exceeds by 20 percentage points. GLM is not the top factor, however – the harness still dictates the winner, and our Louie.ai harness / Opus combo far outperforms opencode / GLM 5.2, 35/59 to 28/59. We’ll see once Louie / GLM 5.2 comes out if GLM has a shot at beating Opus when paired with the best harness.
We also need to send two major warnings. First, on contaminated public benchmarks, we see both Opus and Sonnet do much better than GLM. Anthropic’s outsized performance disappears when run on CyBT-CTF, whose tasks and answers are hidden from model makers to prevent cheating.
The second warning is that our measurements suggest GLM 5.2 may be an illegal distillation of both GPT-5.5 and Opus 4.8. This would help explain how the Goblin Paladin is so close to the presiding champions. Measuring MCC and Cohen’s Kappa scores tells how correlated the right and wrong answers are between models, where 1 means exactly correlated. While OpenAI vs Anthropic have a Cohen’s Kappa of only 0.63, GLM 5.2 jumps to 0.80 and 0.76, respectively. Anthropic reported several months ago that Chinese-origin model companies are performing distillation attacks to steal their model weights, so high correlation scores and similar incorrect answers are noteworthy. Hat tip to Isaac Evans (Semgrep, Founder/CEO) for suggesting we dig into this one.
Let’s break down the bigger findings a bit further:
- GLM 5.2 matches Opus on quality. The model frontier is jagged, and for investigations, GLM 5.2 is a real contender. On CyBT-CTF, GLM 5.2 gets the same solve rate as Anthropic Opus 4.7/4.8 for CyBT-CTF, irrespective of whether Opus is running on the proprietary Claude Code harness or OpenCode harness. Your task set may be different, such as using different tools or needing more instructability, so we encourage using CyBT-CTF numbers only as a starting point.
- GLM 5.2 defeats Sonnet. OpenCode / Sonnet 4.5 is 23/59, vs. GLM 5.2 at 28/59. Compared to Sonnet, GLM 5.2 means paying less and getting more results. Sonnet 4.5 was a breakthrough in the Sonnet series as it was the first time we could trust it with long agentic coding sessions, so GLM 5.2 beating it is impressive.
- The harness choice and prompting setup matters much more than the choice of Opus vs GLM 5.2 . By switching to Louie / Opus 4.x, scores jump from 28/59 to 35/59, so +12%, which is a lot in practice. We’ll be reporting Louie / GLM 5.2 numbers as we get them.
- The one dimension Opus is beating GLM 5.2 on is speed… but probably only for a few weeks/months. Today, Opus 4.8 is 19% faster. However, when Cerebras opens GLM 5.2 support, the current 2X+ cost difference between Opus and GLM 5.2 gives Cerebras a wide margin for beating Anthropic on cost vs speed without compromising on quality.
- The other open models are too far away. While GLM 5.2 competes with the frontier models at 28/59, we find MiniMax 2.5 at 16/59, and for Western models, OpenAI’s GPT-open-120B is 12/59. Such a score difference separates chatting with an AI from delegating tasks to an agent. The release of GLM 5.2 marks the first time we have started to feel comfortable recommending using an open weights model for a frontier-like experience.
We’ll be adding more to the Botsbench pages on GLM 5.2 , Fable / Mythos, and overall board in the coming weeks, so follow our account and site for that. Our results are also a good reminder that owning your AI strategy starts with owning your evals, and the need to be deeply cynical about vendor evals. Likewise, if you are interested in advancing your team and your own journey on using agentic AI for investigations, including how to approach building agents with evals, we welcome you to join our summer Agentic AI x Security training cohorts (Americas + EU/Asia), which builds on the top training of last year’s Black Hat that we helped put together.
Notes from the trenches, monthly.
Graph analytics, GPU viz, and the occasional war story. No spam.
Thanks — we'll be in touch.
