agents
critique
bearish
Agent-generated code, chip designs, and proofs face a fundamental verification challenge: passing 70% of tests is not the same as being correct, and current agents show roughly 0% success on ProgramBench (rebuild a program from its tests)
Machine Learning Street Talk27 Jul 2026