Decomposing Reasoning Efficiency in Large Language Models
Summary of Decomposing Reasoning Efficiency in Large Language Models
We decompose LLM reasoning token-efficiency into truncation robustness, conditional correctness, and workload-/trace-quality-normalized verbosity, to show that efficiency rankings can diverge from accuracy while revealing distinct sources of wasted tokens (verbosity, looping, or logic errors).