Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
arXiv:2509.25420v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved significant advances in reasoning tasks. A key approach is tree-based search with verifiers, which expand candidate reasoning paths and use reward models to guide pruning and selection. Although effective…
