Summary
- ENISA has gained access to Anthropic’s Mythos 5 and OpenAI’s GPT-6 Astra and is testing the models.
- The agency’s access gives European authorities an opportunity to assess advanced cyber capabilities beyond company-supplied descriptions.
- Neither the Commission nor the developers have disclosed the configuration, scope, or results of the current evaluations.
Europe’s cybersecurity agency has begun testing two advanced artificial-intelligence models with significant cyber capabilities, giving a public authority direct access to systems whose ability to find and exploit software weaknesses has become increasingly difficult to assess through ordinary product documentation.
ENISA has obtained access to Anthropic’s Mythos 5 and OpenAI’s GPT-6 Astra, the European Commission confirmed on 10 September. Mythos testing is already under way, while the agency has also been permitted to test OpenAI’s latest model.
The Commission has not disclosed which model configurations ENISA can use, what tools are attached, how long access will continue, or when results may be published. Those details can materially affect cybersecurity performance because a model operating through a restricted interface may behave differently from one connected to code execution, network tools, or a specialist testing environment.
Even with those gaps, the access moves part of frontier-model scrutiny beyond companies evaluating their own systems. Cyber capability has become an important test of whether public institutions can independently examine advanced models rather than relying mainly on benchmarks and risk assessments supplied by their developers.
Cyber capability is unusually difficult to govern
AI systems able to search large codebases, identify vulnerabilities, develop proof-of-concept exploits, or automate portions of security testing have obvious defensive value. The same capabilities can lower the cost of offensive work, particularly when a user can combine a model with tools that scan, execute code, or interact with vulnerable infrastructure.
That dual use makes access politically sensitive. Governments have an interest in understanding what systems can do before the capability becomes broadly available, while developers also have reasons to restrict models whose most powerful cyber functions could be abused.
ENISA’s role creates a third route between those positions. A European agency can test capability without requiring unrestricted public access, although the value of the exercise depends on whether its evaluators see a sufficiently representative version of the model.
Reuters reported that ENISA gained Mythos 5 access after negotiations with Anthropic, while the Commission also confirmed access to GPT-6 Astra. There is no public evidence that the Commission compelled either company to provide these particular models under the EU AI Act, so the current testing should not be described as an enforcement action.
That distinction matters because the AI Act separately gives the Commission formal powers to request access to general-purpose AI models for specified evaluations. The existence of those powers changes the regulatory backdrop, but voluntary or negotiated access remains possible and should not be conflated with compulsory inspection.
Access is only the beginning of evaluation
A regulator possessing credentials for a model still needs a testing environment capable of producing meaningful evidence. Cyber evaluation can involve deliberately vulnerable systems, realistic networks, difficult coding tasks, repeatable benchmarks, and safeguards that prevent an exercise from affecting infrastructure outside the laboratory.
As model capability improves, the expertise required to design those evaluations also becomes more specialised. A weak benchmark can underestimate a model because the task is badly structured, while a permissive laboratory environment can overstate how easily the same capability would translate into uncontrolled real-world use.
ENISA already coordinates European cybersecurity work across incident response, preparedness, regulation, and technical standards, and its recent strategy places greater emphasis on operational preparedness. Testing frontier AI extends that work into an area where much of the evaluation infrastructure still sits with the companies producing the models.
For European oversight to become genuinely independent, agencies will need their own evaluators, compute resources, cyber ranges, secure environments, and relationships with external researchers. Legal authority can open a model, but it cannot substitute for the technical skill required to determine what the model can do.
Published methods will matter
The current tests will be difficult to interpret until ENISA explains its methodology. A headline finding that a model succeeded or failed on a security benchmark says little without knowing whether the model could run code, make repeated attempts, browse documentation, or interact with other tools.
Transparency can be constrained where publishing too much detail would reveal offensive capability, yet meaningful oversight still requires enough information for policymakers and researchers to distinguish an independent assessment from another opaque score.
There is also a wider institutional question. Frontier-model developers iterate quickly, while public evaluation processes traditionally move more slowly. A testing regime that takes months to negotiate access can find itself assessing a model that has already been superseded.
That creates pressure for standing relationships and repeatable procedures rather than one-off negotiations each time a new system appears. ENISA’s access to models from two different developers offers an opportunity to build some of that capability, provided the agency can test comparable functions and retain enough independence to challenge the developers’ own safety claims.
Europe spent much of the AI Act debate deciding which powers regulators should possess. Direct model access shifts attention towards a more practical question: whether public authorities can turn those powers and relationships into technically credible evaluations before the next generation of systems changes the benchmark again.












