Search and filter through extracted claims from AI researchers.
Showing 121-140 of 637 claims in topic "safety"
There is an absence of shared accountability metrics for LLMs across the field
"absence of shared accountability metrics"
"against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate, however expressive it is, because the evidence it would need never appears in the window it sees"
"the obvious repair of carrying a geometrically decaying risk score is insufficient, because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon $N$"
"Under mediated commits and an arbiter detection floor $δ_M$, it bounds the expected number of unauthorized irreversible actions by $B+m-1+m/δ_M$, a constant in $N$, of which the $B+m-1$ term is decided by a model-free rule and therefore survives a fully colluding verifier"
"Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations."
"a monitor retaining cross-iteration state separates the two perfectly"
"when tool outputs no longer merely provide data but begin to specify concrete actions, they effectively become ``commands'' that can drive real-world side effects beyond user intent"
"this risk arises from conflating action induction with execution authorization"
"SARA limits ASR to no more than \(0.63\%\) across four primary evaluation settings while maintaining competitive task utility"
"consistently reduces ASR across additional Agent backbones"
"we propose SARA, which treats action induction and execution authorization as distinct runtime roles and separates action provenance from execution authority"
The EU AI Act's high-risk obligations have applied since August 2, 2026
"the EU AI Act, whose high-risk obligations have applied since 2 August 2026"
No surveyed accountability instrument resolves five identified structural tensions in LLM governance
"five structural tensions that no surveyed instrument resolves"
"it is a synthesis of the surveyed evidence rather than a validated artefact"
"Privacy-Enhancing Technologies (PETs) in computer vision often rely on noise or image perturbations to protect visual data while securely processing it, creating a trade-off between task performance and protection."
"This trade-off is commonly evaluated using image classification, which primarily captures semantic separability and remains robust despite significant geometric, spatial layout or local boundary alterations. As a result, it is too simplistic as a proxy for generic vision tasks."
PETs with similar classification accuracy can differ substantially on other vision tasks
"Across irreversible privacy transformations, key-based block primitives, and learnable image encryption schemes, we demonstrate that PETs with similar classification accuracy can differ substantially on other tasks."
PET evaluation protocols need to move beyond classification-only reporting
"The outcomes highlight the need for PET evaluation protocols that move beyond classification-only reporting."
"In recent years, image-to-video (I2V) generation models have made remarkable progress in subject consistency and temporal coherence, enabling high quality video synthesis."
"However, these advances also introduce new safety risks. Existing studies mainly focus on jailbreak attacks involving single frame violations, while largely overlooking the temporal dimension unique to video generation models."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,951 pending.