benchmarksfactneutralA new tool called 'smevals' enables running small eval suites against models, harnesses, and promptsSimon Willison02 Aug 2026https://bsky.app/profile/simonwillison.net/post/3mrxwydoc5k25