Search and filter through extracted claims from AI researchers.
Showing 161-180 of 637 claims in topic "safety"
Models exhibit behavior that tracks grading authorities rather than users, labs, or law
"behavior that tracks grading authorities rather than users, labs, or law"
Even criminals with very limited skills will be able to target victims at every scale due to AI
"Even criminals with very limited skills will be able to target victims at every scale"
"The importance of having civil society, and not just government and tech companies, centrally involved in devising that plan"
"METR found Sol cheating at the highest rate of any publicly tested model, exploiting test-environment bugs, extracting hidden solutions and covering its tracks."
OpenAI paused reinforcement learning for two weeks post-incident
"disclosing a two-week post-incident pause on reinforcement learning"
"Days later Amazon researchers jailbroke those guardrails, and CEO Andy Jassy told Treasury Secretary Scott Bessent they had used Fable 5 to obtain information usable in cyberattacks, the Wall Street Journal reported."
"No testers have yet been able to find a universal jailbreak"
"AI is reshaping cybersecurity for attackers and defenders alike"
OpenAI is actively working to strengthen its security defenses in response to AI-enabled threats
"OpenAI is strengthening its defenses"
OpenAI is implementing new safeguards that will control the pace of frontier AI model development
"new safeguards are guiding the pace of model development"
OpenAI is strengthening monitoring, alignment, and security measures for frontier AI models
"OpenAI is strengthening monitoring, alignment, and security for frontier AI models"
"ChatGPT for Teens helps teens learn, think critically, and use AI with confidence, with stronger built-in protections, healthy-use features, and additional controls for parents."
ChatGPT for Teens helps teenagers develop critical thinking skills while using AI
"ChatGPT for Teens helps teens learn, think critically, and use AI with confidence"
OpenAI maintains zero data retention policy for eligible API customers
"OpenAI reaffirms Zero Data Retention for eligible API customers"
"previews Private Safety Processing for advanced AI safety without compromising data privacy"
"OpenAI paused improving their models for two weeks after the Hugging Face incident and signs that Astra may cross its critical-cyber threshold."
"Adam Gleave of FAR.AI argues that agent-orchestrated attacks and agentic defenses are already forcing humans out of the loop"
Current monitoring systems have missed the failures they were designed to catch
"current monitoring has missed the failures it was meant to catch"
"highly bio-capable open-weight releases pose a different kind of irreversible risk"
There is a significant gap between lab-internal AI systems and what is available for public access
"how wide the gap is between lab-internal systems and public access"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,951 pending.