OpenAI has announced Aardvark, an “agentic security researcher” built on GPT-5 that continuously scans software projects for exploitable bugs and proposes fixes. The tool launched in private beta on October 30 and is pitched as a way to help engineering and security teams catch issues earlier and at scale.
Aardvark connects to source-code repositories (currently via GitHub Cloud) and builds a threat model for each project. It watches commits, flags likely vulnerabilities, tries to reproduce them in a sandbox to confirm exploitability, and then attaches a suggested patch—generated with OpenAI’s Codex—for human review. OpenAI says the system relies on LLM-driven reasoning and tool use rather than traditional fuzzing or SCA pipelines. (openai.com)
OpenAI reports Aardvark has been running on its own codebases and with external alpha partners “for several months,” surfacing meaningful issues. In internal benchmark tests on “golden” repositories, the agent identified 92% of known and synthetically introduced vulnerabilities, and OpenAI says its work on open-source projects has already yielded multiple CVE-assigned findings.
The private beta is open to a limited set of partners. Requirements include GitHub Cloud integration and a commitment to provide feedback; OpenAI states it will not train its models on participants’ code during the beta.
With software supply-chain risk and LLM-era attack surfaces expanding, Aardvark adds another heavyweight entrant to the growing field of AI security agents. Early coverage from security and tech outlets highlights the agent’s human-like workflow and GPT-5 underpinnings, while noting that broader, independent evaluation will be needed once access widens.
Discover more from MultiMedia
Subscribe to get the latest posts sent to your email.
Leave a Reply