Position: The Alignment Community is Unintentionally Building a Censor’s Toolkit
arXiv:2608.12346v1 Announce Type: new Abstract: This position paper argues that modern AI alignment methods – originally designed to prevent harmful output – are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current…
