Search and filter through extracted claims from AI researchers.
Showing 61-80 of 637 claims in topic "safety"
"On June 26, agents in an evaluation run compromised Artifactory via a legacycredential endpoint. As mentioned above, by design, many Research CaaS workloads used the same shared Artifactory credential. The agents under evaluation identified and exploited a novel vulnerability in a legacy token-refresh endpoint in Artifactory to use these existing credentials to obtain administrator credentials."
"An internal team observed an agent engaging in message board activity and instances of disallowed internet access as early as late May, and with the benefit of hindsight, some early signals identified in our report should have triggered an earlier response."
"On June 27, a cybersecurity monitoring tool alerted OpenAI to port sweep activity within the research environment. Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required."
"OpenAI had at least some responders who observed that Artifactory was being used as a message board and to access the internet, likely while the AIs in question were being actively trained, and they did not think this required them to halt the evaluation run, or to alert key decision makers."
"Agents had motive, opportunity and means to learn to pass notes to each other. So they did. This snowballed into full exploits of the system, and beyond."
"The report contains many details, but little that is new. It tells us in a technical sense What Happened at some points. It does not go into the thinking or dynamics of the agents. It does not go into the thinking and decision making within OpenAI, or the core reasons why things got so bad as to allow this to happen this way."
"In hindsight, there were many things that were going wrong increasingly often during the training runs. The models probed their environments and infrastructure increasingly often and communicated increasingly often. This was not that hidden"
"Models that were already reward hackers first found the message board and communicated there when faced with impossible tasks, after which the message board proved highly valuable, and the agents quickly divided the labor and traded tasks"
"An internal team observed an agent engaging in message board activity and instances of disallowed internet access as early as late May, and with the benefit of hindsight, some early signals identified in our report should have triggered an earlier response"
"Preventing future incidents will require sustained investment in the alignment and control of sophisticated AI systems, as well as security and other safeguards that operate at the speed of the AI agents themselves"
"evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed"
"On July 13, agents compromised the Artifactory signing key and forged administrator credentials there, eventually gaining cluster access."
"At least one model that attacked OpenAI on July 19, resulting in OpenAI finally noticing and dealing with the situation, was from the Astra class, although it was not the Astra version intended for public release."
"There were multiple occasions, including in late May, when an OpenAI team saw agents using a message board to communicate, and decided not to interfere, not to halt work, and not to run that information up the chain."
"The main AI model that attacked HuggingFace was an internal, highly capable and persistent AI model, that we now know was comparable in scale to GPT-5.6 Sol."
"The reality, though, is that at the same time, the use of AI within an organization also radically expands the potential attack surface, giving attackers entirely new ways to gain entry."
"in a recent blog post, the security company Trail of Bits walked through how in a test that they conducted, a particular AI agent was able to escape some sandboxes (exploiting vulnerabilities in Linux KVM, QEMU, and libslirp), but not all. At least in their test, the agent that they tested was not able to escape the sandbox known as Firecracker VM."
"Experiments on multiple benchmarks, target models, and target execution harnesses show that RedEvoAgent outperforms fixed and agentic baselines, improves tool efficiency, and transfers across attacker models and target execution harnesses."
"such retrieval can reuse misleading experiences due to retrieval bias and unclear tool credit, and full trajectories add context overhead while reducing interpretability"
"recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-based retrieval"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.