source: Hugging Face Blog: Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

level: research

researchers from multiverse computing studied how to make language models refuse only harmful parts of a topic, not the entire topic. they focused on political prompts, where a model should answer factual questions but refuse manipulative persuasion. standard safety tuning treats harm as a property of a whole topic, which leads to over-refusal on safe prompts. the team trained models to follow a narrow boundary between harmful and benign prompts within the same topic.

they found that self-generated safety data drops 19.88% of harmful prompts because the model fails to produce a refusal. an escalating retry strategy reduced that to 0.20%. adding benign prompts that look dangerous, like 11,955 examples across 18 types, cut over-refusal on a safe test set from 74% to 5.2%. boundary pairs, where prompts differ only in intent, reduced false refusals on benign prompts from 32.94% to 4.16% while keeping harmful refusal at 87.72%.

the work shows that reporting only harmful refusal rates hides over-refusal problems. a model can look safer by refusing more, but that makes it useless on legitimate prompts. the key is to measure both sides of the boundary and adjust training data composition. this matters for deployments like education or public sector, where a model must answer factual questions but refuse manipulation. the approach can extend to other topics beyond politics.

why it matters: for ai safety, this shows that measuring only harmful refusal rates can hide over-refusal, so developers must evaluate both sides of a topic boundary to build useful models.


source: Hugging Face Blog: Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic