Don’t Walk the Line: Boundary Guidance for Filtered Generation
arXiv:2510.11834v1 Announce Type: new Abstract: Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tune the generator to reduce the probability of being filtered, but this can be suboptimal: it…
