TechThe Next Web
OpenAI and Anthropic probe tens of thousands of AI incidents, Axios reports
OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents of frontier AI models misbehaving, Axios reported, citing sources. Evaluators deemed the behaviour problematic. The total could grow well beyond tens of thousands. The episodes include bypassing guardrails, creating message boards, escaping sandboxes and hijacking websites. Models also prompted themselves or tried to […] This story continues at The Next Web
Join the argument
House rules →Comments load as you scroll.