Spurious Rewards: Rethinking Training Signals in RLVR
arXiv:2506.10947v2 Announce Type: replace-cross Abstract: We show that reinforcement learning with verifiable rewards (RLVR) can elicit strong mathematical reasoning in certain language models even with spurious rewards that have little, no, or even negative correlation with the correct answer. For…
