新闻 · arXiv cs.LG
Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal
Behavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directions in activation space. This raises questions about how such representations are structured, localized, and accessed by optimization. We study adversarial suffix attacks as a probe of representational alignment. We introduce Activation-Guided GCG, which replaces output-based objectives with losses that directly…
en
