Search and filter through extracted claims from AI researchers.
Showing 41-60 of 587 claims in topic "agents"
"How do we build an AI ecosystem where agents, tools, and systems can work together at scale?"
Open standards and projects like MCP, A2A, and Goose are shaping the agentic AI future
"the open standards and projects shaping the agentic future, including MCP, A2A, Goose, etc."
Errors accumulate without effective correction in MLLM agents during extended urban exploration
"errors accumulate without effective correction"
"Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction"
"Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities"
Skill evolution complements model scaling as an approach to improving agent capabilities
"skill evolution complements model scaling"
Persistent knowledge accumulation in the wiki is critical for effective skill evolution
"persistent knowledge accumulation in the wiki is critical for effective skill evolution"
"evolved skills transfer effectively across models and model families, and skills evolved by other models can outperform self-evolved skills"
"larger models generally benefit more from evolved skills, while smaller models with skills can outperform substantially larger models without them"
"Across diverse benchmarks and models, WikiSkill consistently outperforms state-of-the-art skill-evolution methods"
"the first stage performs trajectory-level screening based on process quality, result quality, and data representativeness, selecting a high-quality and representative subset of successful trajectories."
"training on the 10% trajectory subset selected by SWE-Prime outperforms training on the full resolved dataset, yielding relative performance gains of up to 12.2% and 24.2%, respectively."
"task success does not guarantee high-quality supervision: successful trajectories may still contain ineffective, redundant, or risky steps. Directly using such trajectories for SFT can introduce noisy supervision and encourage models to imitate undesirable problem-solving behaviors."
"our in-depth error analysis dissects the distinct drivers of false positives and false negatives, revealing critical weaknesses such as cross-round temporal misalignment and inadequate long-range memory"
"LLMs' performance varies substantially across different defect types and severity levels, with semantically complex or low-salience defects being significantly more likely to be missed"
"experiments reveal that mainstream LLMs exhibit limited overall performance in defect detection and defect lifecycle state tracking, with performance degrading significantly as the number of interaction rounds increases"
"Although recent work explores large language models (LLMs) for automated code review, most approaches oversimplify code review into a single-round, static decision task, which fails to capture the multi-round interactive nature and the complex problem-solving processes inherent in realistic review scenarios"
"The pattern applies when multi-user deployment, execution audit, and expected persona churn hold jointly."
"A probe of a recovered pre-separation build found the governed execution path decoupled from the persona by omission, not by construction; a later wiring change could reverse that isolation, which PES makes an audited architectural rule."
"A development/pilot case in a regulated digital-employee platform records five decisions over one month, each with a rejected alternative. A mechanism check on the shipped implementation found no execution-side re-validation under persona perturbation (five model configurations) and no persona fingerprint on hard-asserted fields."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.