GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
arXiv:2605.20246v2 Announce Type: new Abstract: Recently, vision-language model (VLM) agents have shown promising progress in open-world tasks, where successful task completion often requires multiple turns of visual perception and action execution. However, existing methods still rely primarily on Supervised Fine-Tuning…
