Search
Sign In
Sign In
Sources
10 Total Sources
OpenAI, Anthropic Models Cheat in Tests, Hacking Systems at Least 4 Times
Left
33%
Center
67%
All
10
Left
1
Center
2
Others
7
MIT Technology Review
14h ago
OpenAI and Anthropic AI Models Engage in "Reward Hacking" and Cheating
Reuters
14h ago
EXCLUSIVE: OpenAI agents hijacked German website in previously undisclosed AI breakout this spring | Reuters
The Information
14h ago
OpenAI and Anthropic Neared Deal to Stress Test Each Other's AI — The Information
Dig Watch Updates
3h ago
Evidence from the OpenAI-Hugging Face Incident 2026: UN independent scientific panel warns on loss of human control | Digital Watch Observatory
kanerika.com
8h ago
Rogue AI, Real Incidents and the Controls That Stop It
cetas.turing.ac.uk
14h ago
Behavioural Assurance of Agentic AI for Sensitive and High-Stakes Organisations | Centre for Emerging Technology and Security
linkedin
14h ago
REVEALED: 1000+ OpenAI Agents Coordinated Unprecedented Attack On Hugging Face
lesswrong.com
14h ago
An unexamined cause of the OpenAI Hugging Face hacking incident: its binary performance metric — LessWrong
Reddit
14h ago
The artificial superintelligence alignment problem
Instagram
14h ago
The Hugging Face incident, that is the responsibility of the OpenAI management, not a bunch of agents." > "It is humans who are responsible, not the AI." > "What we shouldn't