agentscritiquebearishCurrent agent evaluations are limited to narrow, verifiable tasks and cannot assess open-ended AI research capabilitiesNarayanan & Kapoor08 Aug 2026https://www.normaltech.ai/p/ai-agents-cant-yet-do-open-ended