Microsoft Research
A 9B model trained on twelve high-fidelity training worlds nearly doubled its base score from 36.5% to 67.1%, coming within fourteen points of GPT-5.4
EvoLib can be applied to any black-box language models and AI systems deployed through APIs because it does not require model updates
EvoLib enables large language models to learn from their own experience during inference without requiring ground-truth labels or external feedback
High simulation fidelity is a must-have for training computer-use agents; shallow worlds hurt the agent while deep ones improve performance
Aeneas allows verifying a large subset of Rust code and provides efficient automation in Lean to support proof efforts for formal verification
AI agents allow scaling automation of formal verification by writing proofs that are independently-verifiable
SymCrypt is releasing verified code, specs, properties, and proofs for SHA-3 and ML-KEM post-quantum cryptography
Aurora 1.5 adds 22 new weather variables relevant to energy, agriculture, transport, and climate risk, plus hourly temporal resolution and probabilistic ensemble forecasting
Aurora 1.5 is released as open source on GitHub with model checkpoints on Hugging Face
Semantic data types help chart compilation systems automatically choose appropriate scales, baselines, formatting, and color schemes
A single Flint specification can compile to multiple charting backends (Vega-Lite, Apache ECharts, Chart.js) without rewriting
Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications
SkillOpt is the best or tied-best method across all 52 evaluation cells covering six benchmarks, seven target models, and three execution modes
AI agents often fail because their instructions are manually modified with no guarantee of improvement
Treating agent skill files as trainable parameters outside frozen models makes agent behavior more reliable without changing model weights
Optimized agent skills transfer across model scales, agent harnesses, and related tasks, suggesting they capture reusable workflow knowledge
To scale agent capabilities, we need a more efficient way to retain and access information over time
Memora achieves state-of-the-art performance on LoCoMo and LongMemEval benchmarks, outperforming Mem0, RAG, and full-context inference while using up to 98% fewer context tokens
Memora dramatically increases agent productivity on long-horizon tasks by decoupling what is stored from how it's retrieved, balancing abstraction and specificity
Current AI agents cannot remember past interactions and must repeatedly be fed relevant information or retrieve it from external sources, which becomes inefficient for long and complex tasks