When Will a Mythos-Level Open Model Arrive?
Hello!
Today we take up the topic that has the AI industry buzzing more than any other lately: "Claude Mythos" and the developments surrounding it.
A month and a half after the announcement, the situation has moved considerably — opposition from the White House, moves by Japan's megabanks, additional AISI evaluations, and a policy shift by Anthropic. So let us pause and answer, with numbers, the simple question: "So when will we actually be able to use the same thing in open source?"
On April 7, 2026, Anthropic announced Claude Mythos Preview.
It is a frontier model said to have reached the top tier of human capability in cybersecurity.
Anthropic offers it as a "gated research preview," limited to the Project Glasswing launch partners (AWS, Apple, Cisco, CrowdStrike, Google, JPMorganChase, Microsoft, NVIDIA, and others) plus more than 40 additional organizations responsible for critical software infrastructure; it is not publicly available (Anthropic official)。
Since the announcement, the industry has kept repeating two questions: "Is Mythos really that strong?" and "When will an open-source equivalent appear?" Public evaluation bodies are beginning to answer the former. As for the latter, following the benchmark numbers reveals a future closer than you might imagine.
In this article, we answer that question with numbers, based on primary sources. Stating the conclusion first:
- an open model matching Mythos on the CyberGym benchmark score arrives in the first half of 2027
(CyberGym is a benchmark measuring how well a model can rediscover known bugs in real software when they are presented as a "problem set"): - an open model matching it on AISI's real-environment benchmarks arrives between late 2027 and early 2028
(AISI is the UK government's AI evaluation body, which runs adversarial tests measuring whether a model can carry a 32-step corporate network attack simulation through from start to finish):
That is the most defensible forecast at this point.
Think of the former as "scores on a fixed problem set"
and the latter as "breakthrough capability in a near-real adversarial environment."
There is a large qualitative gap between the two, and the latter is by far the harder metric. Note, however, that this scenario is based on evaluation data as of May 2026; as discussed below, if AISI's latest observation (task horizons doubling every 4.7 months) continues, the forecast itself could well move earlier.
What exactly was Mythos?
First, the numbers. The benchmark results Anthropic published alongside the Project Glasswing announcement are as follows (compared with Opus 4.6, from the official Anthropic page).
| Benchmark | Mythos Preview | Opus 4.6 |
|---|---|---|
| CyberGym (vulnerability reproduction) | 83.1% | 66.6% |
| SWE-bench Verified | 93.9% | 80.8% |
| SWE-bench Pro | 77.8% | 53.4% |
| Terminal-Bench 2.0 | 82.0% | 65.4% |
| GPQA Diamond | 94.6% | 91.3% |
| Humanity's Last Exam (with tools) | 64.7% | 53.1% |
The numbers cover a broad range, but the standouts are cybersecurity and agentic coding.
According to Anthropic's Frontier Red Team blog, Mythos discovered a vulnerability in OpenBSD that had gone unnoticed for 27 years, and in FFmpeg it found a 16-year-old H.264-related vulnerability that had been missed despite years of fuzzing and human review.
On the Linux kernel, it autonomously chained multiple vulnerabilities all the way to privilege escalation.
What matters here is that these claims have been verified not only in Anthropic's own announcements but by third-party institutions. The UK's AI Security Institute (AISI) is a government-affiliated evaluation body that is becoming the international de facto standard for frontier-model capability evaluation. In its April Mythos Preview evaluation, AISI reported the following results.
- Expert-level CTF tasks: 68.6%
- The 32-step corporate network attack simulation
「The Last Ones」 (roughly 20 hours' work for a human expert): completed 3 of 10 attempts
Mythos was the first model since evaluations began to complete "The Last Ones."
And here comes the important additional information we most wanted to cover in this article.
On May 13, AISI published a new report presenting evaluation results for a newer Mythos Preview checkpoint.
- The Last Ones: Completed 6 of 10 attempts(double the previous 3)
- The other real-environment benchmark, "Cooling Tower": completed 3 of 10 attempts(no model had ever solved it before)
In other words, in just one month, Mythos's real-environment attack-chain capability has grown substantially.
In the same report, AISI notes that "the time horizon of cyber tasks (the task length a model can complete autonomously) is doubling every 4.7 months" and that "this is an acceleration from the 8-month estimate as of November 2025." Mythos Preview and GPT-5.5 overshot even this existing trend.
It was not actually a one-model race
From here, we get to the part that connects directly to this article's title.
In May, AISI tested OpenAI's GPT-5.5 under the same evaluation framework (AISI GPT-5.5 evaluation). The results are as follows.
| Evaluation item | Mythos Preview (initial) | Mythos Preview (new checkpoint) | GPT-5.5 |
|---|---|---|---|
| Expert-level CTF tasks | 68.6% (±8.7%) | - | 71.4% (±8.0%) |
| The Last Ones completion | 3/10 | 6/10 | 3/10 |
| Cooling Tower completion | 0/10 | 3/10 | Not achieved |
In AISI's evaluation, an early checkpoint of OpenAI's GPT-5.5 line also demonstrated cyber capabilities close to the level of Mythos Preview.
As AISI itself has clearly commented, this looks like a
「result suggesting not one company's outlier performance, but rising capability across frontier models as a whole」
.
There is, however, a point that must not be misread. What AISI evaluated was an early checkpoint of GPT-5.5, and AISI itself explicitly states that public deployments carry additional safeguards, monitoring, and access controls, so evaluation results do not necessarily represent the capability available to ordinary users. This is not a claim that a general user can access the same capability through ChatGPT today.
An even more striking example is a reverse-engineering challenge AISI presented: analyzing a custom virtual machine written in Rust and its bytecode, estimated at roughly 12 hours for a human expert. GPT-5.5 solved it in 10 minutes 22 seconds, at an API cost of $1.73. Even allowing for access controls, it is a fact that we have entered an era in which frontier models can be hundreds of times more efficient than human experts on specific tasks.
The reality of "containment" — politics and operations
Before continuing with the technical discussion, we need to take stock of the political and operational developments around Mythos. Operating a "contain and distribute" strategy has entered a phase more difficult than one might imagine.
Opposition from the White House Wall Street Journal reported in late April 2026, and Bloomberg also confirmed, that Anthropic planned to expand Mythos access beyond the initial roughly 50 organizations by adding about 70 more. The Trump administration has voiced opposition. Two reasons are reported: one is concern over misuse; the other is concern that Anthropic lacks the compute resources to support 70 additional organizations, which would impede use by the US government (including the NSA). Anthropic denies the latter.
Early reports of unauthorized access On April 21, Bloomberg reported that unauthorized users gathering on a private Discord forum had obtained access to Mythos Preview on launch day. Anthropic officially commented that it is investigating reports of unauthorized access through a third-party vendor environment. This does not establish that model weights or the system itself leaked, but it illustrates how difficult it is to operate limited distribution of a powerful AI model. According to Fortune, by the time roughly 40 organizations had access, the number of people with access had reached the thousands, and industry experts commented that "a leak was only a matter of time."
Developments in Japan
Developments in Japan The Nikkei reported on May 13, with ITmedia, Yahoo! News, SB Creative's "Business+IT," Ledge.ai, and other specialist outlets following suit, that Japan's three megabanks — MUFG Bank, Sumitomo Mitsui Banking Corporation, and Mizuho Bank — are expected to obtain access rights to Mythos as early as within May.
This will be the first adoption in actual business operations by Japanese companies. Access rights were a main agenda item at the May 12 meeting between visiting US Treasury Secretary Bessent and executives of Japanese financial institutions. On the same day, Prime Minister Sanae Takaichi instructed ministers to examine cyberattack countermeasures, and Finance Minister Satsuki Katayama announced the establishment of a joint public-private task force. On May 14, a working group hosted by the Financial Services Agency was reportedly held, attended by the three megabanks, the Bank of Japan, the AI Safety Institute, and the Japanese subsidiaries of Anthropic, OpenAI, and Google.
The May 18 policy change And in the most recent news — as Reuters reported on May 18 — Anthropic has changed course, adopting a policy that allows Project Glasswing participants to share vulnerability information discovered with Mythos with organizations outside the program (other companies' security teams, industry bodies, regulators, government agencies, open-source maintainers, the media, and the public). Anthropic's communications team explains that "as the program has matured, we have enabled broader sharing of information to maximize its defensive impact."
Put these four developments side by side and it becomes clear that Mythos is no longer merely a technology whose distribution one company manages — it is already becoming a geopolitical resource entangled with finance, government, and critical infrastructure. The limited-distribution policy remains in place, but its operation has begun shifting from the original closed framework toward a gradually more open information-sharing model.
How far along are open models?
Now to the heart of the matter. Consider GLM-5.1, which Z.ai announced on April 7 (official Z.ai documentation、Hugging Face repository)。
This is a roughly 750B-class MoE (Mixture of Experts) architecture open-weight model distributed under the MIT license. The weights are published on Hugging Face, and it can be served locally with SGLang, vLLM, Transformers, and others.
GLM-5.1 benchmark results (Z.ai published figures, based on the official model card):
| Benchmark | GLM-5.1 (Open) | Claude Opus 4.6 | GPT-5.4 | Mythos Preview |
|---|---|---|---|---|
| CyberGym | 68.7% | 66.6% | 66.3% | 83.1% |
| SWE-Bench Pro | 58.4% | (as of previous release) | (as of previous release) | 77.8% |
| Terminal-Bench 2.0 | 63.5% | - | - | 82.0% |
| MCP-Atlas | 71.8% | 73.8% | - | - |
Note that the figures above are, at this point, primarily Z.ai's own published values, and should be read separately from full replication by independent institutions.
The number to watch is CyberGym.An open-weight model has surpassed Claude Opus 4.6 and GPT-5.4. The gap to Mythos is 14.4 points. That may look large at first glance, but tracking the trajectory of CyberGym scores tells a different story.
The previous generation, GLM-5, scored 48.3 on CyberGym. GLM-5.1 scores 68.7 — a 20-point rise in a single release cycle. To quote Z.ai CEO Lou: "At the end of last year, agents could only execute about 20 consecutive steps. GLM-5.1 can run 1,700 steps. That is the difference four months makes."
Of course, there is no guarantee that per-release gains continue at this rate. But extrapolating mechanically on CyberGym alone, a scenario in which the generation after GLM-5.1 approaches Mythos's 83.1% is entirely plausible. Given Z.ai's historical release cadence (four to six months), we consider late 2026 to the first half of 2027 a reasonable window for closing the gap on CyberGym numbers.
"Chinese models are mostly distillation" no longer holds
Here we would like to address a misconception that persists stubbornly in the industry.
"DeepSeek and Qwen are surely just distilling the outputs of OpenAI and Anthropic."
That is the view.
Around 2024 it was partially true; in 2026 it is off the mark.
The reason is that
original technology from Chinese labs has become reference material for the international research and implementation community
.
Multi-head Latent Attention(MLA) is an attention mechanism first proposed in DeepSeek-V2 (DeepSeek-V2 paper). By compressing the KV cache into low-rank latent vectors, it cuts memory usage by 93.3% while maintaining modeling quality. Independent researchers have verified that it is clearly superior to earlier KV-cache reduction methods such as GQA. The prominent US researcher Sebastian Raschka described MLA in his LLM architecture commentary as the idea that defines the DeepSeek era.
Symbolic of this is the title of the paper "Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs". It is the work of a research group from Fudan University and Shanghai AI Lab proposing a method (MHA2MLA) for retrofitting MLA onto existing MHA-based models such as Llama. A design originating from a Chinese lab is being referenced as the standard toward which existing mainstream architectures are being modified.
DeepSeekMoE, auxiliary-loss-free load balancing, Multi-Token Prediction — these too are DeepSeek originals, and have become standard pieces of today's open-source MoE model design. GLM-5.1's roughly 750B scale stands on the MoE recipe DeepSeek pioneered.
Distillation has not disappeared entirely. It is still used for small derivative models (the DeepSeek-R1-Distill line and the like). But for flagship pretraining, there are areas where original research has become the international reference base.
What this means is that the technical foundations for building a "Mythos-equivalent open model" are largely coming together on the Chinese-lab side as well. The remaining variables are compute budget, data, and building evaluation environments specialized for cybersecurity.
Two capability dimensions, two arrival dates
The phrase "Mythos-level" is, frankly, imprecise. Mythos's capabilities should be viewed along at least two dimensions, each with its own arrival date.
Dimension 1: benchmark score level (CyberGym / SWE-Bench Pro)
This one is relatively near. Reaching CyberGym 83.1 is within range of the next generation if you mechanically extrapolate the growth rate of the current open-source leader, GLM-5.1. On SWE-Bench Pro as well, GLM-5.1 posts 58.4 with Z.ai's internal implementation — a 19.4-point gap to Mythos's 77.8.Between late 2026 and Q1 2027, we consider it a fully viable scenario that an open-weight model reaches Mythos-equivalent territory on CyberGym.
Dimension 2: AISI real-environment benchmark level (completing The Last Ones and Cooling Tower)
This is a substantially harder domain. Chaining a 32-step corporate network attack from start to finish requires not benchmark optimization but a composite of skills: long-horizon agency, tool use, memory management, and state tracking.
Moreover, in AISI's latest May evaluation, the new Mythos checkpoint stabilized The Last Ones at 6 of 10 completions and broke through Cooling Tower (previously unachieved by any model) for the first time. In other words, the very "Mythos level" the open side must catch up to shifted upward substantially in just one month.
GLM-5.1's documentation advertises "up to 8 hours, 1,700 steps of autonomous execution," but that is under controlled conditions — different from an adversarial environment like AISI's. For the open side to reach this point, late 2027 to early 2028 seems the reasonable estimate. That said, if the "doubling every 4.7 months" pace AISI observes continues, this timeline too could move earlier.
Dimension 3: operational capability in the real world
This is a different matter — completing benchmarks and real-world operation are not the same. As AISI itself stresses in its report, the evaluation environment has constraints: no active defenders, no defensive tooling running, no retaliation for tripping alerts. How models behave in a real world where defenders also wield AI has not yet been evaluated by anyone. AISI has announced it is developing new evaluation environments that include active defense, so the true touchstone is still to come.
Why weights get released — or don't
Here we must touch on an uncomfortable question. Will Z.ai really keep releasing models with GLM-5.1-class capabilities as open weights?
There are both optimistic and pessimistic readings.
The optimistic case
Z.ai, DeepSeek, and Qwen have all fundamentally maintained open-weight strategies. This is a commercial strategy, but it also has a national-strategy dimension: protecting China's AI ecosystem as a whole from a Western proprietary monopoly. Heightened cybersecurity-specific capability is not, so far, a strong reason to change course abruptly.
The pessimistic case
That said, once capabilities transferable to offense cross a certain threshold, the possibility that they fall under Chinese government export controls is not zero. A movement symmetrical to the US treating Mythos-class models as a national-security matter could well occur.
One more reference point is Anthropic's own official roadmap. Anthropic explains that while Mythos Preview will not be made generally available, it is building out the safeguards needed to eventually deploy Mythos-class models safely. On the open-weight side too, if cyber-specific capability crosses a certain level, Z.ai and others may decide to separate standard models from limited-distribution models. That is the two-track scenario: the standard GLM-5.2 stays open, while GLM-5.2-Cyber goes to limited distribution.
In that case,
a "Mythos-level open model" may be reached in capability terms yet never be publicly released as a distribution
.
That would be a political impossibility, not a technical one.
The answer to the opening question
Let us give the opening question the most honest answer the available primary sources allow.
An open-weight model that is "Mythos-equivalent" on CyberGym benchmark numbers: Q4 2026 to Q1 2027
The leading candidate is Z.ai's GLM-5.2 line, followed by DeepSeek V4/V5 and Qwen3.6 onward. Even if the GLM-5 line's growth rate (20 points per release) halves to 10, the next release lands at 78–79.
An open model that is "Mythos-equivalent" at the AISI real-environment level (completing TLO / Cooling Tower): Q3 2027 to Q2 2028
This is structurally unfavorable terrain for the open side: evaluation-environment buildout, adversarial simulation, long-horizon agent capability. Expect a lag of about a year. However, if the "doubling every 4.7 months" pace AISI observes continues, this timeline could compress further.
Three caveats apply, however
First, reaching a capability and deciding to release it are separate questions. It is entirely possible that Z.ai or DeepSeek decides to switch to limited distribution precisely because Mythos-equivalent capability has been reached. In that case, a benchmark-equivalent open model would exist in capability terms but not be reflected in the publicly released version.
Second, Mythos itself is not standing still. As AISI's May report showed, in one month Mythos improved from TLO 3/10 to 6/10 and Cooling Tower 0/10 to 3/10. Anthropic has laid out a roadmap of maturing safety mechanisms in the next Opus line and deploying Mythos-equivalent capability generally. The "frontier level" that Mythos defines may itself have moved elsewhere by the time the open side catches up. This is a chasing game, and the goalposts move.
Third, as Anthropic's May 18 policy shift shows, the Glasswing model itself has begun moving from closed enclosure to controlled information sharing. Now that vulnerability findings from Mythos can be shared with organizations outside the program, regulators, and the media, the boundary between "until an open version arrives" and "until closed-side knowledge becomes common" is becoming far blurrier than originally assumed.
What this means for business practitioners
Let us close by organizing what this shift means from the perspective of CTOs, CISOs, product owners, and business leaders weighing AI adoption.
First,
the premise that cybersecurity is "a specialist domain for organizations that have specialists" is collapsing
. In an era when AI can do 12 hours of expert work in 10 minutes for $1.73, if your security stack amounts to "human experts look at it periodically," you may be unable to keep pace with AI-accelerated vulnerability discovery, verification, and remediation.
One caveat is worth adding here.An analysis reported by Reuters on May 20 indicates that among security practitioners, the immediate post-announcement narrative that hacking capability was about to be unleashed all at once is viewed by some as overstated. Semgrep CEO Isaac Evans commented that "there is a major perception gap between practitioners and policymakers" and that Mythos, while certainly a technical advance, has drawn reactions about its real-world function that are not grounded in reality. In other words, improved discovery capability, real as it is, does not immediately mean attacks increase by orders of magnitude. The practical bottleneck lies less in discovery than in the subsequent verification, prioritization, remediation, and rollout. Building AI-assisted vulnerability scanning into business workflows cannot be postponed, but investment decisions driven by excessive fear can backfire. AISI reports that cyber task capability is doubling every 4.7 months, and next term's security planning needs to be designed with that pace as its premise.
Second,
the strategic value of open-source LLMs is changing.
Through 2025, a simple routing rule sufficed: open for cost savings, closed API when quality matters. In 2026–2027, when specific capabilities (cyber, for example) become restricted on the closed side, situations arise in which the open side is the only option. The reverse also holds: you may judge an open model's capability too high and refrain from using it. Technology selection comes to include axes beyond cost.
Third,
the crude classification "Chinese models = national-security risk / distilled knockoffs" should be discarded。
. Technologies that are now standard in open-source LLM design — MLA, DeepSeekMoE — originated in Chinese labs. Unless risk assessment and respect are treated as separate matters, technical decisions will go wrong.
Fourth — and this is the most important:
not a posture of waiting for "when will it come out," but
a posture of designing "what to do when it does."
What matters is not guessing when a Mythos-equivalent model will be released. It is reviewing your software assets, dependency libraries, patching processes, AI usage logs, permission management, and auditability on the premise that both attackers and defenders will use ever-stronger AI. And, as the Reuters May 20 analysis showed, the practical bottleneck lies not in discovery but in the verification, prioritization, remediation, and rollout that follow. This is precisely why the operational design companies need is a pipeline that does not stop at "the AI found it" — that is, one that runs discovery → verification → prioritization → remediation → audit logging → permission design as a single integrated loop.
The advance of AI capability cannot be stopped. Which is exactly why what companies need is not to fear the models, but to design operations, security, and governance as one. Anthropic's May 18 relaxation of information-sharing restrictions within Glasswing is also an opportunity for companies. The UK's NCSC is positioned to make use of Mythos-derived vulnerability information, and in Japan such information will begin circulating through the FSA's working group. Building — now — the pipeline that turns received information into executable countermeasures is the survival strategy for next year and beyond.
About Qualiteg's hands-on consulting
Questions like the one this article addressed —
"how do we connect the evolution of frontier AI capability to our own operations, security, and governance?" — are continuous design work, not a one-time decision. The considerations span model selection, risk assessment, building operational flows, ensuring auditability, and internal organization.
Qualiteg maintains a dedicated consulting team that works through these challenges together with client companies and stays alongside them. We can support each phase, from technical evaluation of AI models to security operations design, governance, and talent development.
"What happens when a Mythos-equivalent open model arrives?"
"Will our current security stack hold up as it is?"
"We want to move forward with AI adoption — where should we start?" We will build concrete answers to questions like these together with you, from both the technology and business sides. Please feel free to reach out.
▶ Contact us here:
https://qualiteg.com/contact?inquiry=consulting_business
See you next time!
Main sources (primary and near-primary):
- Anthropic official: Project Glasswing、Mythos Preview Frontier Red Team blog
- UK AI Security Institute: Claude Mythos Preview evaluation (April)、GPT-5.5 evaluation (May)、The pace of progress in autonomous cyber capability (May 13)
- Bloomberg: White House Opposes Anthropic's Mythos AI Expansion (April 30)、Mythos Unauthorized Access (April 21)
- Wall Street Journal: White House Opposes Anthropic's Plan to Expand Access to Mythos Model
- Reuters: Anthropic to let partners share Mythos cybersecurity findings (May 18)、Fears of unfettered hacking spurred by Anthropic's Mythos AI model overstated (May 20)
- The Nikkei: Three megabanks including MUFG Bank to obtain access rights to the Mythos AI (May 13, 2026)
- ITmedia, SB Creative's "Business+IT," Ledge.ai: coverage of the three megabanks and the Takaichi administration's response
- CyberGym paper (arXiv 2506.02548) and leaderboard
- Z.ai GLM-5.1: official documentation、Hugging Face repository, and independent commentary by MarkTechPost, Awesome Agents, and Lushbinary
- Sebastian Raschka, "LLM Architecture Gallery: MLA"
- DeepSeek-V2 paper (arXiv 2405.04434)、Fudan University MHA2MLA paper (arXiv 2502.14837)