benchmarksopinionneutralCurrent evaluation methods fail to capture 'big model smell' that differentiates top modelsswyx & Alessio29 Jul 2026https://www.latent.space/p/ainews-claude-opus-5-fable-level