Summary
- OpenAI says preliminary evaluations of an upcoming model called Astra mean it cannot rule out the company’s Critical cybersecurity capability threshold.
- That threshold covers autonomous discovery and exploitation of zero-day vulnerabilities in hardened systems or execution of novel end-to-end attacks.
- Development, testing, monitoring, network access, and model security controls are being tightened before OpenAI decides how such capability could be deployed.
OpenAI has tightened security around an upcoming AI model after internal testing suggested its cybersecurity capabilities may be strong enough to cross a threshold the company associates with materially new forms of cyber risk.
The company said on 7 August that preliminary evaluations of the model, known internally as Astra, showed substantial advances in agentic coding and cybersecurity. OpenAI has not concluded that Astra has reached its Critical capability level, but said performance was strong enough that it could no longer rule out that classification while testing continues.
Under OpenAI’s Preparedness Framework, the Critical cybersecurity threshold covers models capable of autonomously finding and developing working zero-day exploits across many hardened real-world critical systems, or devising and executing new end-to-end strategies for attacking hardened targets from only a high-level objective.
That represents a considerably higher bar than making programmers more productive or helping security teams investigate known vulnerabilities. OpenAI’s disclosure indicates that testing is approaching the point where a model could remove far more of the specialist human work required to conduct sophisticated cyber operations.
Security controls move into model development
OpenAI has responded by imposing tighter controls before the capability assessment is complete, including isolated testing environments, restricted access to networks and tools, stronger protection and encryption for model weights, expanded monitoring, and sandboxed execution. Internal Astra activity that does not meet those requirements is being paused rather than allowed to continue under controls designed for less capable systems.
The company is also applying monitoring across agentic uses of Astra during training and evaluation, with systems intended to identify risky actions or signs of misalignment and trigger human security review. Government agencies and selected AI safety organisations are expected to participate in further capability testing, while third-party evaluators will receive additional guidance on the controls needed for higher-risk workloads.
Those measures show how frontier-model security is beginning to resemble the protection of other sensitive technology assets, where access to systems, code, credentials, networks, and test environments is constrained according to the damage misuse could cause. A laboratory evaluating ordinary coding capability can tolerate a different threat profile from one testing whether an agent can independently discover vulnerabilities and act against hardened targets.
The distinction becomes harder as evaluation itself grows more realistic. Testing advanced cyber capability requires sufficiently demanding environments, tools, and objectives to determine what models can actually do, yet the same realism creates risk if the environment is poorly isolated or model actions reach systems that were never intended to form part of the exercise.
Capability creates a deployment problem
A model reaching the Critical threshold would not automatically make widespread autonomous cyberattacks inevitable because practical harm still depends on access, safeguards, tools, targets, credentials, and deployment design. The classification would nevertheless change the assumptions providers and customers have to make when an AI agent is connected to development infrastructure or given permission to interact with external systems.
Enterprise security architecture consequently moves beyond questions about what employees are allowed to ask a model. Agentic systems can receive credentials, network connectivity, development tools, browser access, APIs, and permission to execute code, so governance increasingly depends on controlling what the system can do after producing an answer rather than merely filtering what it says.
Permission design becomes part of model risk management under those conditions. An organisation may want an AI agent capable of finding weaknesses across its own estate, but the usefulness of that capability rises alongside the consequences of excessive privileges, compromised credentials, prompt injection, configuration errors, or a model acting outside its intended scope.
Security teams will therefore have to treat powerful agents more like privileged automation than ordinary productivity software, with segmented environments, narrow permissions, audit trails, independent monitoring, and explicit limits around external connectivity. Model providers can impose controls at the service layer, but customers still determine much of the surrounding infrastructure into which those systems are placed.
Cyber defence sits on the same capability curve
The policy problem is complicated by the fact that capabilities creating offensive risk can also strengthen defence. Models able to inspect large codebases, reproduce vulnerabilities, generate patches, and automate security testing could help defenders address weaknesses faster, particularly in organisations that lack enough specialist staff to review everything manually.
OpenAI says it wants advanced cyber models to help defenders identify and fix vulnerabilities before attackers exploit them. Restricting capability too heavily could therefore withhold useful tools from legitimate security teams, while broad availability before safeguards mature could reduce the skill and effort required for harmful activity.
That tension becomes harder to manage as models improve because capability cannot be divided neatly into offensive and defensive categories. Understanding how to exploit a vulnerability can be necessary when verifying that a patch works, while the same reasoning can support an attacker attempting to compromise an unpatched system.
OpenAI’s disclosure is therefore more consequential as a security governance development than as a preview of another model release. Astra remains under evaluation and the company has not said it definitively crosses the Critical threshold, but development controls are already being changed on the assumption that it might.
If subsequent testing confirms that assessment, organisations deploying the next generation of cyber-capable agents will have to make access controls, isolation, monitoring, and permission boundaries part of the procurement decision itself. The model may supply the capability, while much of the practical risk will be determined by the systems and authority placed around it.












