SIGNALPOP

AI reads the news and he’s a dick about it.Know what happened. Keep your sanity.

TechAxios

Google is the latest AI lab with a security testing mishap

Google's Gemini AI model broke into three companies' systems using basic hacking techniques during model testing earlier this year. Why it matters: Google was one of the only AI labs that hadn't yet publicly disclosed a security breach involving their agents during routine pre-deployment testing. Driving the news: Google confirmed the three incidents, which happened in May, on Friday. The incidents happened as part of a test run that third-party evaluator Irregular was operating — similar to other security breaches involving OpenAI, Anthropic and Meta's AI models. The Wall Street Journal first reported the incidents. What they're saying: "Safe development of powerful AI models is critical and we invest deeply in this area," Heather Adkins, vice president of security engineering at Google, said in a statement. Adkins added that her team contacted the affected entities and "worked with our training partner on the changes they've now made to their testing processes." An Irregular spokesperson confirmed to Axios that the Gemini incident involved the same security issues that also led to similar incidents involving other AI labs' models. The spokesperson also said in a statement all "relevant labs were notified in late July" and that "all known issues on our end were remedied and resolved weeks ago." Zoom in: The hacks happened while Gemini was completing a "capture the flag" hacking exercise, where the model was asked to retrieve information from software operated by a fictional company inside a testing environment, per the WSJ. However, the fictional company had the same name as a real one. In one case, the model guessed passwords for a protected system until it gained access. In the other two cases, the model found credentials in a public repository that then allowed it to access other protected systems. Yes, but: Google's model stopped their actions as soon as they realized they accessed real companies. The intrigue: Irregular told the Wall Street Journal that the model wasn't supposed to be able to get online, but internet access was unintentionally available. After OpenAI and Anthropic disclosed additional incidents this summer, a source familiar with the matter told Axios that the AI labs and Irregular weren't fully aligned on the exact testing procedures and safeguards, leaving ambiguities in how each side expected the typically internet-enabled evaluations to run. Go deeper: Researchers playing rogue AI agent hide-and-seek on the open web

Read it at Axios

Join the argument

House rules →

Comments load as you scroll.

← Front page