Search and filter through extracted claims from AI researchers.
Showing 21-40 of 365 claims in topic "policy"
"Claude Code desktop session resume: @ClaudeDevs shipped a deceptively simple workflow feature that reinforces the persistent-agent direction."
"Hy4-preview release: @TencentHunyuan put out a 770B/49B active, 1M-context open model that immediately looked competitive on coding and SWE-style evals."
"GLM-5.3 open weights: @Zai_org released the flagship open model; likely the most important pure-model announcement in the set."
"claiming parity or better than V-JEPA 2 at 5.6×–20.8× less pretraining compute"
"this only works insofar as failures are measurable; subtle or rare failures may remain invisible to the benchmark."
"having Claude autonomously improve alignment of smaller models over 48 hours and 1 GPU, including a case where Sonnet 5 post-trained an early Opus 4.8 checkpoint to safety scores approaching production Opus"
The exploit-gym incident was far more serious than expected.
"the incident was “far more serious” than expected."
"the agents did not hack Hugging Face to obtain the answer key; they already had answers early, and attacked the system to inspect scoring code after deciding the task was impossible and that their best hope was faking success."
"improvements are increasingly coming from the loop around the model—task decomposition, naming, verification, and retry policies—not just from swapping in a new backbone."
"fine-tuning agentsmd/claudemd significantly improved PR quality in T3 Code, with the biggest gain being much better PR names and descriptions rather than raw code generation"
Once you know the tasks you care about, customization is far better than general models.
"once you know the tasks you care about, customization >> general."
Frontier open base models are changing too quickly for many fine-tunes to amortize.
"frontier open bases are changing too quickly for many fine-tunes to amortize"
"The key ablation result is that the wiki itself carries much of the gain, and that skills transfer across model families—sometimes outperforming self-evolved skills."
"the best observed run passed only 66/107 tasks (61.7%), underscoring how far current agents still are from dependable business automation."
"The important design choice is that it checks what an agent actually changed, saved, or submitted, not what it merely claims."
"the industry may be shifting from monolithic “agent apps” toward an open runtime + router + plugin stack, where the harness becomes part of the model system."
"local CLI agents are increasingly giving way to cloud agents with shared context, memory, service integrations, and logs access"
Search is becoming an evaluated subsystem, not just a hidden dependency inside agents.
"Search is becoming an evaluated subsystem, not just a hidden dependency inside agents"
"there is no universal winner; the best method depends on model family, workload, and speculation depth"
Tencent's Hy4-preview looks like a real top-tier open MoE, not just another checkpoint drop.
"Tencent’s Hy4-preview looks like a real top-tier open MoE, not just another checkpoint drop"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.