An idea judged on paper is judged on its pitch. Execution is where it meets reality, and plenty of ideas that looked strong on paper don’t survive it.
Si, Hashimoto, and Yang tested this with research ideas. Reviewers rated LLM-generated ideas more novel than ideas from human experts. Then experts spent over 100 hours executing each idea, and after execution the LLM ideas lost much of their score, enough that the ranking flipped.
So evaluating strategy on paper overstates what models can do, and in general we should be wary of flawless ideas. It’s also an argument for testing ideas against the territory rather than the map, which gets easier as AI moves validation downstream.
References
- Chenglei Si, Tatsunori Hashimoto, and Diyi Yang, “The Ideation–Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas” (opens in new tab), 2025