SABER: Small Actions, Big Errors – Safeguarding Mutating Steps in LLM Agents
arXiv:2512.07850v1 Announce Type: new Abstract: Despite rapid progress in LLM agents, performance on long-horizon, tool-using tasks remains fragile. To better understand this fragility, we ask a simple question: emph{do all actions contribute equally to failure?} Analyzing execution traces on $tau$-Bench…
