AI DevelopmentAnalysis5 min readPublished September 29, 2026

A capability report with a specific test population

Anthropic: Open GLM-5.3 Nearly Matches Mythos at Exploits

In Anthropic's tests the open GLM-5.3 model succeeded at 50 of 410 exploit attempts, close to Claude Mythos Preview at 56. What the report shows and omits.

DA
Digital Applied Team
Research and practical guidance
CoverageSeptember 29, 2026

Anthropic reports that the open-weight GLM-5.3 model approached Claude Mythos Preview on one exploit-development benchmark. The practical implication is that defenders should plan around capable models being broadly accessible. The result needs its denominator and testing conditions attached: it is not proof that the models are equivalent across cybersecurity work.

Editorial note: Prepared October 1 as a September 29, 2026 dispatch, using the dated announcements cited below. Later product developments are outside this article’s scope.

Key takeaways
  1. 01
    Preserve the denominatorThe headline result concerns attempts on a particular benchmark, not all vulnerabilities.
  2. 02
    Separate capability from safeguardsA model’s ability to solve a task and its willingness to comply are different properties.
  3. 03
    Prioritize defensive readinessPatch exposure, access controls and response drills remain actionable without reproducing attacks.

01 — The evidenceWhat the benchmark numbers establish

In its September 29 report, Anthropic records 50 successful end-to-end exploit attempts out of 410 for GLM-5.3, versus 56 for Claude Mythos Preview on ExploitBench. That is approximately 12.2% and 13.7%, calculated from the published counts. Anthropic tested a competitor’s model; these are vendor-run findings. They should not be presented as an independent certification.

Source: Anthropic, September 29, 2026. Shares are Digital Applied calculations from the reported counts; they describe this benchmark only.
ModelSuccessful attemptsAttemptsCalculated share
GLM-5.35041012.2%
Claude Mythos Preview5641013.7%

An attempt is not the same unit as a unique vulnerability. Repeated trials can explore different paths against a target. Dividing successes by attempts answers a narrower question than asking how many distinct systems are exploitable. Keep that unit visible in a slide or purchasing memo so the number does not grow into a claim the experiment never made.

The six-attempt difference also does not by itself establish statistical equivalence. That conclusion would need a defined analysis of uncertainty and the trial structure. The useful reading is that both models demonstrated the relevant capability in this evaluation.

02 — Practical implicationsThree questions that should stay separate

Anthropic additionally reports safeguard bypass rates of 64–100% across its simulated tests. Its altered-weight experiment used about 2,200 GPU hours and reduced refusals from above 90% to 2–12%. These are separate experiments with different denominators from ExploitBench. The report says tested Claude comparisons used safeguards disabled where noted; do not silently substitute the behavior of an ordinary public endpoint.

Capability
Can the model complete the task?
Evaluation result

Read the target set, tools and success definition before comparing scores.

Technical ability
Behavior
Will it follow the request?
Safeguard result

Examine the tested configuration and the categories of request used.

Deployment condition
Access
Who can run that configuration?
Distribution policy

Distinguish downloadable weights from a restricted hosted service.

Operational exposure

A high refusal rate in a public endpoint is not evidence that every deployment of the same weights behaves identically. Conversely, a laboratory configuration designed to study capability does not describe every customer-facing product. Procurement and security teams should record the exact model artifact, serving configuration and available tools in their own evaluation.

Defensive scope

This article summarizes capability and safeguard findings. It does not provide exploit procedures or instructions for removing safety controls.

03 — Practical implicationsUse the independent assessment for context

The report points to the September 17 CAISI assessment, which places GLM-5.3 roughly four months behind the US frontier on its aggregate cyber benchmarks. That is CAISI’s comparison, with its own methodology. It is not a forecast that every open model will close every gap on a four-month schedule.

A model comparison becomes more useful when the reader can follow each measure back to its source. Do not combine one organization’s capability score with another organization’s refusal rate to create a new overall ranking. The tasks, scoring rules and configurations would first need to align. The same caution applies to claims about cost: a reported demonstration budget is not the expected cost of every future attack.

Our cyber-model access guide separates distribution restrictions from capability claims. Access conditions matter because they influence how widely a tested configuration can be used, but restrictions alone do not measure the quality of defensive work.

04 — Practical implicationsTurn the report into a defensive work list

Start with exposure you can actually reduce. Identify internet-facing services, confirm patch ownership and check that emergency updates can reach the systems that need them. A written patch policy is incomplete if an old deployment has no owner or a change cannot be rolled back safely. Choose one ordinary update and trace it from advisory to verified installation.

Digital Applied defensive recommendations; these are not additional findings from Anthropic’s experiment.
WorkstreamConcrete evidence
PatchingAn owner, current inventory and a verified deployment record.
Agent accessA documented tool and credential boundary for each workload.
DetectionA tested alert that identifies the affected system and task.
ResponseA drill showing the operator can revoke access and stop work.

For internal AI testing, use owned systems and isolated targets with explicit authorization. The question can be whether a defensive assistant correctly explains a patch or triages an alert; it does not require recreating an exploit. Keep evaluation data separate from production credentials and preserve the results needed for review.

The DNS containment incident illustrates why tool boundaries and a tested stop path deserve attention even in research environments. Our permission-default guide applies the same principle to business agents.

05 — Practical implicationsMake the response proportional to the evidence

The report supports a stronger assumption of broadly available cyber capability. It does not identify which of your systems is vulnerable, prove a particular attacker is using the model or replace a threat assessment. Use it to prioritize a concrete review of exposure and response speed, then measure whether that work improves the organization’s position.

Our AI transformation service can help define controlled evaluations and deployment permissions for defensive AI workflows. Keep the acceptance criteria focused on accurate, reviewable work and a clear boundary around the systems the agent may touch.

Next step

Reduce exposure while preserving the test limits

Treat the findings as evidence that useful and harmful cyber capabilities are spreading. Keep the benchmark conditions attached to the comparison, and put the operational effort into patching, scoped access and a response process that has been exercised.

Agentic AI implementation

Build a workflow you can evaluate and control

Digital Applied helps teams connect AI capabilities to useful work, clear acceptance checks and responsible operating limits.

Task evaluationsCost visibilityControlled access
Start with one task

Define the pilot

  • →Approved source material
  • →A named reviewer
  • →A clear acceptance check
  • →Spending and permission limits
Questions and answers

Practical questions

No. The headline counts concern one exploit benchmark, and the report includes other experiments with separate conditions.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source