Search and filter through extracted claims from AI researchers.
Showing 81-100 of 637 claims in topic "safety"
"LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone."
"ModelAudit produced definitive security decisions for all 135 families (100%)"
"Conditional on making a definitive judgment, ModelScan achieved 100% precision, recall, and F1."
"ModelScan for 67 (49.6%)"
Fickling provided no unique detection value beyond the combination of ModelAudit and ModelScan
"Fickling identified no unique true- positive families beyond those found by the combination of ModelAudit and ModelScan."
"These findings underscore the need to separate judgment accuracy from judgment availability, as well as incremental detection coverage from tool-level redundancy."
"for the 48 malicious families where ModelScan failed to complete its analysis, both ModelAudit and Fickling generated detections consistent with ground truth."
"AI-generated text detection is commonly framed as a binary document-level judgment about whether a text is human-written or machine-generated. This framing breaks down for mixed-origin writing, where content origin and expression origin may differ."
"our disclosed D2C-Routing-based detector system reaches 0.8603 four-way Avg TPR@1%FPR, 6.5 points above the same-split RACE-local rerun"
"error analysis shows that distinguishing AI-content/human-expression from fully AI-generated text remains the hardest boundary"
"verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored ("the user is testing my ability to follow instructions"), safety-flavored ("the user is testing my boundaries"), both, or neither: framings that predict compliance very differently"
"On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across all tested steering conditions"
"A CoT-prefill intervention on eval-awareness-negative rollouts suggests the link is causal, with 10 of 11 prefills shifting compliance in the predicted direction"
"eval-awareness is not behaviorally uniform: aggregate suppression rates can move while the safety-relevant component does not"
"the same "X% suppression of eval-awareness" can correspond to qualitatively different behavioral outcomes"
"Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed"
"As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions."
"Using chain-of-thought (CoT) monitoring, we find that harmful execution is often preceded by intent signals in reasoning."
Post-hoc chain-of-thought labels are too coarse to show how intent changes during generation
"However, post-hoc CoT labels are too coarse to show how intent changes during generation."
"The probability of calling an intent tool provides a judge-free, fine-grained signal of the model's tendency to pursue that behavior."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,951 pending.