Search and filter through extracted claims from AI researchers.
Showing 141-160 of 637 claims in topic "safety"
"we investigate three attack scenarios and uncover a temporal vulnerability in I2V systems: unsafe semantics may emerge not from a single frame, but from semantic composition over time."
"Experiments on closed-source commercial models, including Kling, Seedance, Veo and PixVerse, show that TempJail improves attack success rate over prior state-of-the-art methods by 23.3\% under GPT-5.2 evaluation and 22.0\% under human evaluation."
Two key challenges in temporal jailbreak attacks are temporal abstraction and semantic camouflage
"We further identify two key challenges in such attacks: temporal abstraction and semantic camouflage."
"Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property."
Black-box access to LLM judges exposes their capabilities to model extraction attacks
"However, black-box access exposes these capabilities to model extraction attacks."
"Existing extraction methods do not specifically target LLM judges and provide limited support for multiple evaluation protocols under restricted query budgets."
"Extensive experiments on state-of-the-art LLM-as-a-judge and reward models show that JUDGESTEALER consistently outperforms existing extraction baselines, achieving up to 73.3%, 87.0%, and 71.6% accuracy for pointwise, pairwise, and listwise evaluation, respectively."
JUDGESTEALER demonstrates robustness against representative extraction defenses
"Moreover, JUDGESTEALER demonstrates robustness against representative extraction defenses."
"JUDGESTEALER exploits the strong cross-protocol agreement to acquire pointwise scores and transform them into pairwise and listwise supervisions without additional victim queries."
Computing Lipschitz constants is difficult even for shallow ReLU networks
"computing them is difficult even for shallow ReLU networks"
"for every fixed $p\in (1,\infty)\cap \mathbb{Q}$, maximizing the $L_p$-norm over a zonotope in $\mathbb{R}^d$ is W[1]-hard with respect to the dimension $d$"
"our hardness results imply that brute-force enumeration algorithms are essentially optimal for this problem under the Exponential Time Hypothesis"
"Our paper resolves an open problem posted at COLT'25"
The authors used LLMs as part of their research process
"we explicitly describe our research process including the use of LLMs"
There is currently no systematic plan for how society will address AI
"The urgent need for—and lack of—a systematic plan for how society will address AI"
AI creates risks around bioterrorism, deepfakes, disinformation, and cyberattacks
"The risks that AI creates, around bioterrorism, deepfakes, disinformation, cyberattacks, and so on"
"reward-seeking can produce motivated reasoning, cleaner-looking but less trustworthy chains of thought, and behavior that tracks grading authorities rather than users, labs, or law"
Models can diagnose when they are in a deception test and still rationalize lying behavior
"models that reason about the grader or safety review board, diagnose a deception test, and still rationalize lying"
"The stakes are whether chain-of-thought monitoring can remain useful as reasoning traces become enormous, compressed, and harder for humans or other models to audit"
"cleaner-looking but less trustworthy chains of thought"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,951 pending.