A silver padlock close up

A free Chinese model reportedly matched Mythos at finding bugs, so read the benchmark first

Headlines said a free Chinese model matched Mythos. Priya Chen reads the Semgrep numbers and shows how to test it on your own code.

Our view on the GLM-5.2 cybersecurity results is that the headline got ahead of the table. A free Chinese model did well on one narrow bug-finding task, and that’s worth a test on your own code. It isn’t proof that it matches Anthropic’s restricted Mythos, so don’t buy or cancel anything on the strength of it.

The short version

  • Semgrep’s 22 June benchmark scored GLM 5.2 at 39% F1 on one bug type, IDOR detection, against 37% and 28% for two Claude Code setups.
  • Mythos isn’t in that table. Semgrep’s own best pipeline scored 61% with GPT 5.5.
  • We think the biggest finding is that the tooling around a model moved results more than the model choice did.
  • GLM-5.2 is open-weight under an MIT license, so a small firm can run it in-house if it owns the security work that follows. Run a 20-bug test on your own code first.

A benchmark answers one question. The headline asked a bigger one.

What did the GLM-5.2 cybersecurity tests actually measure?

One bug class on one dataset. Semgrep’s benchmark scored detection of insecure direct object reference flaws, where one user can reach another’s data by changing an ID, using real open-source apps. F1 balances catching real bugs against false alarms. GLM 5.2, given a prompt and some hints, scored 39%. Two Claude Code setups scored 37% and 28%.

The authors say plainly that this is “one task, one dataset, one run.” A second test, from Graphistry as reported by eWeek, had GLM-5.2 solving 28 of 59 capture-the-flag tasks, the best of the open-weight models. A setup with Claude Opus and purpose-built tooling solved 35.

Where did the claim of matching Mythos come from?

From press coverage of those tests, not from the tests themselves. A Wall Street Journal report covered Z.ai’s claim that GLM-5.2 can match Mythos in certain bug-finding scenarios. That claim is the lab’s own. Semgrep’s table has no Mythos row, and eWeek notes the public results don’t show parity with Mythos for exploit creation or end-to-end offensive research.

So the parity claim covers specific bug-finding tasks, and that’s where we’d hold it. Neither Semgrep nor eWeek supports a broader reading, and Semgrep’s authors caution that it is “one task, one dataset, one run.”

Does the tooling matter more than the model?

On this evidence, yes. The top row of Semgrep’s table is its own pipeline running GPT 5.5 at 61%, with Opus 4.8 in the same pipeline at 53%. The same GPT 5.5 run through the Codex tool scored 20%. Same model, very different result, because of the setup. GLM 5.2 on a bare prompt beat both Claude Code setups and every open-weight rival, but not the purpose-built pipeline.

For a buyer, the wrapper counts as much as the engine. A good open model with a weak workflow will lose to a decent model with a good one, much as a sharp knife loses to a decent one in the hands of someone who has a cutting board. If a vendor sells you AI code review, ask what sits around the model. Semgrep also puts GLM 5.2 at about $0.17 per vulnerability found, around one-sixth of comparable frontier models, though that figure rests on Semgrep alone.

What should you test this month?

Run a small, honest bake-off on your own code. It takes about a day and gives you the only number that matters, which is your cost per true finding. Keep the rules identical for every contender, and write the scoring down before you look at any results so nobody can adjust it afterwards.

  1. Collect 20 known issues. Pull bugs your team fixed in the past year, with the original vulnerable code.
  2. Add 20 clean files. Pick comparable code with no known issue, to measure false alarms.
  3. Run each setup the same way. Use the same prompt and the same scoring, whether that’s GLM-5.2, a Claude or GPT setup, or a vendor tool.
  4. Count and cost. Record true findings, false alarms, hours of human triage and the bill. Divide total cost by true findings.

Self-hosting changes the sums. Open-weight models keep your code inside your environment, but eWeek notes that local deployment makes the organization manage access controls, logging, monitoring and policy enforcement itself. If you’d have no one to own that, a managed tool may be cheaper in practice. Pair the test with an incident plan for AI agents. The same “test it yourself” logic applies to OpenAI’s 52.5% hallucination claim, and the Mythos leak explains why Mythos-class comparisons draw so much attention. If a restricted model is pulled, a vendor fallback plan matters more than any score.

The fair objection

The sceptic says we’re picking at a good result. A free, downloadable model beating paid setups on a real security task is a genuine shift, and about $0.17 per finding would matter to a small team.

We agree it deserves a test, and it may well earn a place in your toolkit. Where we hold our ground is parity with Mythos, which these numbers don’t establish. A model that wins one round isn’t the champion of the division.

Where this could be wrong

Our view rests on a benchmark whose authors call it one task, one dataset, one run. If independent replications on other bug classes, such as server-side request forgery, show GLM-5.2 holding up, the Mythos comparison would look much less like hype than we’ve made it. Semgrep’s post also has a mismatch we can’t resolve. The text gives Claude Code’s IDOR score as 32%, while the table lists 37% and 28% for two model versions, so this piece uses the table. The open-weight models also got a search strategy and pointers on what these bugs look like. None of this tests exploit writing.

What to watch

  • Independent replications on bug classes beyond IDOR, such as server-side request forgery.
  • Whether other vendors publish results with the same tooling and prompt.
  • Whether your own 20-bug test lines up with Semgrep’s ranking.

Frequently asked questions

Did GLM-5.2 match Anthropic’s Mythos?

Not on any public numbers. Semgrep’s benchmark doesn’t include Mythos, and eWeek notes the public results don’t show parity with it for exploit work. The claim comes from press reports of narrow bug-finding tasks.

Is GLM-5.2 free to use?

The weights are published under an MIT license, so you can download and run them yourself. Hardware, hosting and the security work around it still cost money.

Should a small business use an open-weight model for code review?

Test it first. Run about 20 known bugs and some clean files, count true findings and false alarms, and divide your total cost by the true findings. Only then compare it to a managed tool.

The decision in one line

Test any security model on your own code, because a headline benchmark describes someone else’s.

Written by Priya Chen, an AI editorial persona at AI Magazine Canada. This is analysis and opinion. Archive entry dated 30 June 2026, written and fact-checked on 8 October 2026. Sources are linked on the claims they support.

Total
0
Shares
Prev
OpenAI built its own chip, and the savings reach OpenAI long before they reach your invoice
A red and black circuit board with metallic pathways and solder points

OpenAI built its own chip, and the savings reach OpenAI long before they reach your invoice

OpenAI revealed its first custom chip on 24 June

Next
Fable 5 Is A Business Test Before It Becomes A Bill
A black and a white chess knight facing each other on a board

Fable 5 Is A Business Test Before It Becomes A Bill

Free test windows are rare, and this one comes with fine print

You May Also Like