Search and filter through extracted claims from AI researchers.
Showing 141-160 of 587 claims in topic "agents"
"RingCentral uses ChatGPT Work and Codex to accelerate AI product development and centralize operational intelligence across engineering and operations"
Enterprises are adopting agentic AI
"enterprises are adopting agentic AI"
Frontier firms are pulling ahead in AI adoption
"frontier firms are pulling ahead in AI adoption"
"Models keep absorbing the harness into their weights — soon, it will be a harness for human attention rather than for the model."
"Large language models process large amounts of information but usually lack an explicit mechanism for maintaining compact and evolving conceptual representations."
"Most LLM-based automated algorithm design methods optimize a designated component within a human-specified scaffold, fixing overall organization and component interactions."
"Across four NP-hard problems, ATLAS outperforms several state-of-the-art component-synthesis methods and a matched full-synthesis baseline while remaining competitive with strong human-designed algorithms."
"Our results suggest that embedding-guided quality-diversity search can make the enlarged full-algorithm design space practically searchable."
Qwen 3.8 27B successfully built a script to transform its own .jsonl transcripts to Markdown
"I set Pi up with Qwen 3.8 27B and had it build a script for transforming its own .jsonl transcripts to Markdown... which it did!"
"On 68 OpenML classification tasks, LACE with GPT-5.4-mini significantly outperforms auto-sklearn, H2O, and a fixed XGBoost baseline, with no detectable difference against AutoGluon, the strongest search-based system evaluated, while covering the full benchmark."
"To our knowledge, LACE is the first to formulate general tabular pipeline AutoML this way, evaluated on standardized OpenML tasks under a leakage-controlled protocol that withholds dataset identity from the generator."
Tabular foundation models are more accurate than LACE but only on the subset of tasks they support
"Newer tabular foundation models are more accurate on the subset of tasks they support, but apply a fixed pretrained predictor rather than returning an editable task-specific program."
"There is a shift happening where the ability to train a base model to be a general agentic reasoner is becoming opaque like at-scale pretraining practices from a few years ago."
"We investigate whether agentic artificial intelligence can automate parts of the process of designing genetic programming systems by introducing an agentic framework that identifies and implements parent selection algorithms using large language model (LLM) reasoning and retrieval-augmented generation."
Agentic AI provides a step toward automated configuration and design of evolutionary systems
"These findings demonstrate the potential of agentic AI to translate domain knowledge into generating executable components, providing a step toward automated configuration and design of evolutionary systems."
"CAD raises mean final best fitness in all eight domain and API comparisons. Across all CAD runs, learned libraries are adopted by most later programs and repeatedly rediscover validation, reachability, and structural utilities. These results support that discovering reusable primitives improves evolutionary program search for content generators."
"Large language models can generate executable programs, which makes it possible to search directly over procedural content generators rather than individual levels."
"The key claim many engineers highlighted is that the capability jump came entirely from scaled post-training/RL on longer-horizon executable tasks, not from a larger base model"
"Harnesses are becoming an optimization target in their own right: A few posts reinforced that benchmark and product gains are increasingly coming from the scaffold/harness layer, not just base-model IQ."
"a companion post argues current frontier agents are much stronger when source is available than when they must reason over binaries"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,951 pending.