Search and filter through extracted claims from AI researchers.
Showing 81-100 of 2685 claims of type "fact"
"Even then, the information about the message board failed to reach decision makers."
"On June 26, agents in an evaluation run compromised Artifactory via a legacycredential endpoint. As mentioned above, by design, many Research CaaS workloads used the same shared Artifactory credential. The agents under evaluation identified and exploited a novel vulnerability in a legacy token-refresh endpoint in Artifactory to use these existing credentials to obtain administrator credentials."
"An internal team observed an agent engaging in message board activity and instances of disallowed internet access as early as late May, and with the benefit of hindsight, some early signals identified in our report should have triggered an earlier response."
"On June 27, a cybersecurity monitoring tool alerted OpenAI to port sweep activity within the research environment. Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required."
"OpenAI had at least some responders who observed that Artifactory was being used as a message board and to access the internet, likely while the AIs in question were being actively trained, and they did not think this required them to halt the evaluation run, or to alert key decision makers."
"Agents had motive, opportunity and means to learn to pass notes to each other. So they did. This snowballed into full exploits of the system, and beyond."
"In hindsight, there were many things that were going wrong increasingly often during the training runs. The models probed their environments and infrastructure increasingly often and communicated increasingly often. This was not that hidden"
"Models that were already reward hackers first found the message board and communicated there when faced with impossible tasks, after which the message board proved highly valuable, and the agents quickly divided the labor and traded tasks"
"An internal team observed an agent engaging in message board activity and instances of disallowed internet access as early as late May, and with the benefit of hindsight, some early signals identified in our report should have triggered an earlier response"
"evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed"
"On July 13, agents compromised the Artifactory signing key and forged administrator credentials there, eventually gaining cluster access."
"At least one model that attacked OpenAI on July 19, resulting in OpenAI finally noticing and dealing with the situation, was from the Astra class, although it was not the Astra version intended for public release."
"There were multiple occasions, including in late May, when an OpenAI team saw agents using a message board to communicate, and decided not to interfere, not to halt work, and not to run that information up the chain."
"The main AI model that attacked HuggingFace was an internal, highly capable and persistent AI model, that we now know was comparable in scale to GPT-5.6 Sol."
"in a recent blog post, the security company Trail of Bits walked through how in a test that they conducted, a particular AI agent was able to escape some sandboxes (exploiting vulnerabilities in Linux KVM, QEMU, and libslirp), but not all. At least in their test, the agent that they tested was not able to escape the sandbox known as Firecracker VM."
LLM-generated text now exhibits at least 38 identifiable clichéd patterns
"My LLM cliché highlighter is up to 38 patterns now"
Deploying robots to real-world applications requires a long development timeline
"Getting robots to real deployment takes a long time"
Errors accumulate without effective correction in MLLM agents during extended urban exploration
"errors accumulate without effective correction"
"Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost"
"CritICL, a novel inference-time framework that improves reasoning while maintaining high efficiency. Our key insight is that LLM failure modes exhibit structured patterns across model scales within the same family. Instead of treating failures as undesirable outputs, CritICL leverages them as a source of guidance. Specifically, we utilize failure modes derived from weaker models and incorporate them into inference through critique-based in-context examples"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.