Search and filter through extracted claims from AI researchers.
Showing 101-120 of 4537 claims
"the agents did not hack Hugging Face to obtain the answer key; they already had answers early, and attacked the system to inspect scoring code after deciding the task was impossible and that their best hope was faking success."
"improvements are increasingly coming from the loop around the model—task decomposition, naming, verification, and retry policies—not just from swapping in a new backbone."
"fine-tuning agentsmd/claudemd significantly improved PR quality in T3 Code, with the biggest gain being much better PR names and descriptions rather than raw code generation"
Once you know the tasks you care about, customization is far better than general models.
"once you know the tasks you care about, customization >> general."
Frontier open base models are changing too quickly for many fine-tunes to amortize.
"frontier open bases are changing too quickly for many fine-tunes to amortize"
"The key ablation result is that the wiki itself carries much of the gain, and that skills transfer across model families—sometimes outperforming self-evolved skills."
"the best observed run passed only 66/107 tasks (61.7%), underscoring how far current agents still are from dependable business automation."
"The important design choice is that it checks what an agent actually changed, saved, or submitted, not what it merely claims."
"the industry may be shifting from monolithic “agent apps” toward an open runtime + router + plugin stack, where the harness becomes part of the model system."
"local CLI agents are increasingly giving way to cloud agents with shared context, memory, service integrations, and logs access"
Search is becoming an evaluated subsystem, not just a hidden dependency inside agents.
"Search is becoming an evaluated subsystem, not just a hidden dependency inside agents"
"there is no universal winner; the best method depends on model family, workload, and speculation depth"
Tencent's Hy4-preview looks like a real top-tier open MoE, not just another checkpoint drop.
"Tencent’s Hy4-preview looks like a real top-tier open MoE, not just another checkpoint drop"
"@Yuchenj_UW reported GLM-5.3-Flash at 270 tok/s, 10% higher quality than GLM-5.2 on OfficeQA Pro v2 at 1/10 the cost"
A 239GB 2-bit variant of GLM-5.3 retains about 81% accuracy after shrinking from 1.51TB.
"@UnslothAI claimed a 239GB 2-bit variant retaining about 81% accuracy after shrinking from 1.51TB"
Z.ai's GLM-5.3 family moved from strong API model to broadly deployable open weights.
"Z.ai’s GLM-5.3 family moved from strong API model to broadly deployable open weights"
Grok Bot is a viable competitor to Codex/ChatGPT.
"Grok Bot is a viable competitor to Codex/ChatGPT"
GPT 5.6 is a serious coding alternative to the Claude 5 series.
"Today, GPT 5.6 is a serious coding alternative to the Claude 5 series"
"these have come a long way from the early on-demand kids books created with traditional coding practices."
Current AI-generated content still has rendering glitches and coherence issues.
"I noted a couple places where the rendering glitched or the model lost a little bit of coherence in the story"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.