I gave ChatGPT / Sol and Claude / Opus the same game design document, then played both builds.
Here's what worked, what broke, and which game I'd rather keep playing. This is my hands-on playtest, not a controlled benchmark.
0:00 Same brief, two models
0:31 ChatGPT / Sol playtest
2:32 Claude / Opus playtest
6:51 The comparison and verdict
7:34 Which game would I keep playing?