When in Doubt, Plan It Out: Committed Small Language Model Deliberation for Reactive Reinforcement Learning
arXiv:2606.16995v1 Announce Type: cross Abstract: Reinforcement Learning (RL) policies often degrade in unfamiliar environments because they lack explicit deliberation. We propose Plan, Align, Commit, Think (PACT), a hybrid architecture that combines a fast, reactive RL policy with a slow, deliberative…
