OpenAI says its AI agents breached its own systems before Hugging Face

Weeks before OpenAI’s agents hacked Hugging Face, the agents worked together to find and exploit a vulnerability in the infrastructure supporting the company’s cybersecurity testing, OpenAI researchers said Wednesday.

The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI.

OpenAI’s internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the company’s testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity conference on Wednesday.

The genius-level actions observed in testing showed exactly what could happen if these models make it to the hands of hackers.
“We believe this is a watershed moment for computer security as an industry, as well as for OpenAI as a whole,” Michael Dalton, a member of OpenAI’s technical staff, said during the presentation.

OpenAI started testing its internal research model, which is not intended for public release, on May 7. Within a day, after getting stuck while trying to complete a cybersecurity evaluation, the model reasoned that it could reach the internet indirectly. The model then discovered it could write its own files in Artifactory’s shared package repository. It then left a note to other agents about its discovery.
Soon, more agents started leaving notes for each other in the repository, creating a de facto message board where the agents collaborated and traded information about their findings, including new vulnerabilities they found.

(Sam Sabin – Axios) 

The Global is in partnership with SIOI

Latest articles

Related articles