Lewis Tunstall
GPT-5.6-luna being cost-effective has caused OpenAI servers to be constantly overloaded
Gemma4 models produce consistently concise outputs when enable_thinking=false, unlike Qwen3.5
Frontier post-training increasingly looks like expert training followed by distillation
5 trillion tokens of high-quality code data has been released on the Hub
Poolside published all trajectories of their evaluations, similar to what Llama 3 did
AI models are now capable of paperclip maximizing behavior in practice