Search and filter through extracted claims from AI researchers.
Showing 61-80 of 587 claims in topic "agents"
"Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning"
"while orientation and pedestrian-aware movement remain unreliable"
"Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited work) traceable. A single trust domain does not satisfy both cheaply."
"We present Persona-Execution Separation (PES): persona and execution reside in different trust domains, connected by a governed contract bridge. The persona is singly-homed and may drift; execution is faceless and audited."
"Under LLM representational indistinguishability, any single-domain mechanism that meets all three must re-introduce typed change objects, an external gate, and a stable audit anchor: PES rebuilt at higher coupling cost."
NPO achieves comparable or better performance than GEPA with fewer rollouts
"NPO achieves comparable or better performance than GEPA with fewer rollouts"
"prompt optimization emerging as a promising approach capable of delivering performance gains comparable to those achieved by fine-tuning model weights, while reducing computational costs in both optimization and serving"
Recent prompt optimizer developments favor unnecessarily complex approaches
"recent developments increasingly favor unnecessarily complex prompt optimizers"
"its advantage increases with stronger teacher models, suggesting that stronger teacher reasoning can partially substitute for optimizer-side search complexity"
NPO-optimized prompts transfer well to other student models, especially within the same model family
"NPO-optimized prompts elicit similar performance improvements when applied verbatim to other student models, especially across models within the same family"
"simple, linear prompt optimization can rival substantially more sophisticated and complex search procedures"
"Efficiently improving autonomous agents across diverse tasks is central to accelerating recursive self-improvement (RSI) in agentic AI"
"The strongest model we test, gpt-5.6-sol, matches or outperforms the best existing method on almost all evaluated instances."
"This holds even at level 2, where the returned algorithm is fixed before seeing the evaluation instances."
"Performance also improves sharply across models released less than eight months apart, suggesting that this capability is moving quickly."
"for the well-specified operations problems we study, a single untuned LLM query can already produce algorithms competitive with specialized methods."
"These results suggest that frontier LLMs can be a serious empirical baseline for algorithm design in well-specified OR problems."
"Existing propose-and-verify methods typically score every candidate on a fixed task set, wasting rollouts on unrelated behaviors and allowing aggregate scores to obscure specific regressions."
"Across three agent harnesses and four benchmarks, HarnessLens improves average held-out performance by 7.6-13.6% while consuming substantially less evaluation budget than competing baselines."
"These results demonstrate that behavior-aware verification with explicit attribution enables more reliable and sample-efficient harness evolution under constrained interaction budgets."
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.