OpenAI Says Astra Crosses Critical Cybersecurity Threshold, Finds 2 Zero-Day Vulnerabilities
Image: 디지털투데이

OpenAI Says Astra Crosses Critical Cybersecurity Threshold, Finds 2 Zero-Day Vulnerabilities

02 September, 2026.Technology and Science.10 sources

Developing · updated 2h ago · 10 outlets

Astra becomes first OpenAI model to cross critical cybersecurity threshold, autonomously finding and exploiting flaws. OpenAI will restrict access to Astra’s advanced cyber capabilities during its rollout.

10 outlets1 divide2 facts unevenly coveredseverity 2/10

Beat 1 · The verdict

WIRED stresses limited partner access; OpenAI focuses on system-card transparency.

Beat 3 · What got skipped

5 Other outlets never mentioned: Astra found and exploited two zero-day vulns in combination

El Economista · Hipertextual · OpenAI · Pluang · citybiz

4 Western Mainstream outlets never mentioned: OpenAI says safeguards would have prevented the Hugging Face incident

CNBC · Fortune · TechCrunch · WIRED

Read them yourself

Do not take our word for it. Here is what they published.

All 10 outlets

Full story

Astra hits “Critical”

OpenAI said its forthcoming AI model Astra is the first model to cross its “Critical” cybersecurity capability threshold, with the company saying it can find previously unknown security flaws and develop ways to exploit them without step-by-step guidance from humans.

Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.

OpenAIOpenAI

In an OpenAI briefing, Amelia Glaise (아멜리아 글레이스) said, “Astra can find unknown security flaws in systems with multiple layers of defenses and develop ways to exploit them, even without human intervention.”

Image from CNBC
CNBCCNBC

OpenAI said it applied additional safeguards and warned that safeguards could wrongly judge normal work as misuse or unauthorized behavior, potentially delaying or halting unrelated tasks and long-running agent tasks.

OpenAI also said that in its testing process Astra found 2 zero-day vulnerabilities and figured out ways to exploit them in combination, and that it strengthened training to more reliably reject harmful cyber requests after the Hugging Face incident.

OpenAI researcher Fouad Martin (푸아드 마틴) said, “It can help defenders find and fix vulnerabilities, but without safeguards it can make attackers stronger,” as the company prepares a controlled rollout of Astra’s advanced cybersecurity functions.

Limited access, new safeguards

OpenAI said it plans to make Astra available “soon,” but access to its most advanced cybersecurity capabilities will be more limited, with advanced cybersecurity work initially restricted to a group of testers.

OpenAI said it will expand defensive use through Daybreak Blue, and it planned to disclose more information about Astra’s safety, security and alignment evaluations in the model’s system card at launch.

Image from El Economista
El EconomistaEl Economista

In its own update, OpenAI said it delayed parts of Astra’s development and release while it strengthened and tested protections against cyber misuse and unauthorized model actions, and it said it incorporated learnings from the Hugging Face incident into its safety approach.

WIRED reported that OpenAI said its misalignment monitor may “occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, leading to it inadvertently being slowed, paused, or stopped.”

WIRED also said OpenAI planned to limit everyday users from accessing Astra’s advanced cyber capabilities, including a new “misalignment monitor,” and that partners in Daybreak would get early access to a less restricted version at launch.

Benchmarks and risks ahead

OpenAI said Astra’s preparedness evaluation combined automated public and private benchmarks with expert-driven assessments, and it described Astra as a significant increase in cybersecurity capabilities compared to GPT‑5.6 Sol.

Astra achieves a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities.

OpenAIOpenAI

In the ExploitBench evaluation, OpenAI said Astra achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities, and it said it discovered and used two zero-day vulnerabilities as part of an exploit chain.

OpenAI said it is in the process of disclosing these two vulnerabilities to the maintainers, and it said Astra results shown reflect capabilities with Daybreak Blue access, not the default production configuration.

The company also described a scenario in which Astra built a full browser-compromise chain that escaped the sandbox and executed commands on the host when the browser opened an HTML file, and it said it also found multiple vulnerabilities in a hardened operating system and combined them into a local privilege-escalation chain from an unprivileged user to root.

OpenAI framed the stakes through its safeguards, saying it needs to cover “Malicious actors using the model” and “The model taking unauthorized, misaligned actions,” as it prepares to release Astra with access controls designed to reduce the risk of severe cyber harm.

SourcesOpenAIOpenAI