Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
arXiv:2510.09694v1 Announce Type: new Abstract: Large models (LMs) are powerful content generators, yet their open-ended nature can also introduce potential risks, such as generating harmful or biased content. Existing guardrails mostly perform post-hoc detection that may expose unsafe content before…
