Explore Briefly, Then Decide: Mitigating LLM Overthinking via Cumulative Entropy Regulation
arXiv:2510.02249v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable reasoning abilities on complex problems using long Chain-of-Thought (CoT) reasoning. However, they often suffer from overthinking, meaning generating unnecessarily lengthy reasoning steps for simpler problems. This issue may…
