MIT and Harvard Teach Language Models to Ask Better Questions, Lifting a Small Model's Battleship Win Rate From 8% to 82%
An ICLR paper from MIT CSAIL and Harvard shows Monte Carlo inference helps Llama 4 Scout outpace GPT-5 at a Battleship test bed for around 1% of its cost.