I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything.
An empirical benchmark across four LLMs reveals that increasing reasoning effort to 'high' rarely improves task accuracy outside complex logic puzzles, while consistently inflating token usage and inference costs.