Three separate AI security incidents involving OpenAI, Anthropic and Meta have been traced to the same Israeli cybersecurity startup, revealing a common link behind disclosures made by the companies over the past two weeks.
The startup, Irregular, hosted the independent evaluation environment used to test advanced AI models before release. During those evaluations, AI models from all three companies accessed websites or systems that were intended to remain inaccessible as part of controlled cybersecurity testing.
OpenAI said a misconfiguration within Irregular’s testing environment allowed ChatGPT to access the public internet. Anthropic later disclosed that its Claude model may also have reached the internet during testing and said it notified Irregular after identifying the issue. Meta subsequently confirmed it had experienced a similar incident and is investigating what happened.
In a statement to CNBC, Irregular said all three incidents resulted from the same evaluation-environment issue first identified during Anthropic’s testing. The company said the events did not involve an AI model escaping its testing environment or carrying out a sophisticated cyberattack. It also stated that there are currently no unresolved issues and that it is preparing a white paper on best practices for secure AI cybersecurity testing.
The incidents have drawn attention to the role of independent firms in evaluating increasingly powerful AI systems. Rather than conducting every assessment internally, major AI developers often rely on third-party specialists to test models for security weaknesses and unintended behaviour before deployment.
According to experts cited by CNBC, independent testing helps ensure objective evaluations by avoiding situations where AI developers effectively assess their own systems. At the same time, the latest disclosures show that as multiple companies depend on shared testing infrastructure, weaknesses in those environments can have consequences across several AI developers. The incidents have therefore shifted attention from the AI models themselves to the security and reliability of the systems used to evaluate them before release.


