Search and filter through extracted claims from AI researchers.
Showing 61-80 of 494 claims of type "critique"
"MM-LLMs inherently lack an understanding of system states and do not track state transitions, often leading to hallucinated actions that deviate from the intended goal"
"generating action plans in natural language tends to limit the generated plans to a high level, introducing ambiguity in action execution"
"most existing neural operator architectures enforce boundary conditions indirectly through training from data even though the boundary condition is often known exactly"
"existing modifications and approaches that do enforce boundary conditions explicitly suffer from impractical restrictions, including boundary smoothness, uniform grids, and separable, box-like domains"
"Yet latent transitions are commonly realized with Transformer-based predictors whose inductive structure is centered on token interaction rather than temporal evolution."
"Existing work spans many agent domains, but domain-centered organization and heterogeneous evaluation often obscure common generation mechanisms and conflate candidate construction with verification and selection."
"Research into event-based object classification methods are hindered by the lack of high-quality vision datasets to use."
"This approach has several practical flaws. The size, weight and power consumption of the device could prohibit deployment at the extreme edge or in covert sensing environments. Besides this, there are security concerns inherent in cloud-based or other off-device computation approaches due to the requirement of sending and receiving potentially sensitive data. Furthermore, this transmission of data introduces latency and requires consistent connectivity to the cloud infrastructure to function."
"Existing reinforcement-learning methods typically output a single reactive action at each timestep, which limits their ability to represent diverse short-term avoidance strategies."
"we observe an evaluation artifact in common crowd-navigation benchmarks: without explicit boundary constraints, learned agents may leave the valid domain and bypass dense crowds"
Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions
"Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions"
"However, existing benchmarks suffer from ''fragmentation'', manifested in limited temporal coverage, limited medium diversity, and incomplete script types."
"existing evaluations often allow shortcuts based on transcripts or single-modality solutions, obscuring whether models genuinely ground predictions in speech"
"their performance on these tasks has not been systematically evaluated under controlled conditions"
"Use-case benchmarks show whether one agent completes one task, but not how changing capabilities, models, runtime mechanisms, capacity, and enterprise data should be owned, changed, admitted, or evidenced together."
"However, no non-taxonomic relations are extracted, highlighting limitations of closed, taxonomy-oriented relation vocabularies"
LLM outputs are treated as authoritative even when they are ungrounded or incorrect
"fluent, confident outputs are treated as authoritative even when ungrounded or incorrect"
Current LLM accountability frameworks suffer from under-specification of human oversight
"Four persistent gaps emerge: under-specification of human oversight, absence of shared accountability metrics, disciplinary disconnection, and limited empirical evaluation"
There is an absence of shared accountability metrics for LLMs across the field
"absence of shared accountability metrics"
"they use static basic function to constructed graph filter which cannot effectively adapt to the frequency-domain distribution of graph data"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.