Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
arXiv:2605.21282v2 Announce Type: replace Abstract: Online off-policy reinforcement learning (RL) is shaped by two coupled choices: the policy class and the update rule. Gaussian policies are fast and have tractable entropy, but struggle with multimodal action distributions. Generative policies are…
