CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
arXiv:2606.11063v1 Announce Type: cross Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untrusted model’s trajectory. If the trusted model detects such an…
