Paper Type

Complete

Abstract

Large language models (LLMs) are increasingly embedded in organizational decision-making, yet current safety evaluations focus almost exclusively on overt harm refusal. This paper identifies a subtler but more pervasive risk: moral rationalization facilitation. This occurs when users seek AI assistance in justifying ethically questionable workplace decisions. We introduce MoralRatBench, a benchmark grounded in neutralization theory (Sykes & Matza, 1957) that evaluates LLM resistance to moral rationalization requests across 12 neutralization techniques, 20 workplace scenarios, and 5 output formats, yielding 1,200 two-turn test cases per model. The benchmark includes 40 control scenarios to ensure models remain helpful for ethical requests. Evaluation of 8 frontier LLMs reveals substantial variation in resistance to moral neutralization. A multi-turn analysis showed that initial refusal does not guarantee sustained resistance. Certain neutralization techniques, particularly those framed in moral language, are systematically harder for models to resist. MoralRatBench will be released as an open-source evaluation tool.

Paper Number

1638

Comments

AI SYSTEM

Share

COinS
Best Paper Nominee badge
Top 25 Paper Badge
 
Aug 15th, 12:00 AM

Beyond Overt Harm: Benchmarking LLM Resistance to Moral Neutralization Requests

Large language models (LLMs) are increasingly embedded in organizational decision-making, yet current safety evaluations focus almost exclusively on overt harm refusal. This paper identifies a subtler but more pervasive risk: moral rationalization facilitation. This occurs when users seek AI assistance in justifying ethically questionable workplace decisions. We introduce MoralRatBench, a benchmark grounded in neutralization theory (Sykes & Matza, 1957) that evaluates LLM resistance to moral rationalization requests across 12 neutralization techniques, 20 workplace scenarios, and 5 output formats, yielding 1,200 two-turn test cases per model. The benchmark includes 40 control scenarios to ensure models remain helpful for ethical requests. Evaluation of 8 frontier LLMs reveals substantial variation in resistance to moral neutralization. A multi-turn analysis showed that initial refusal does not guarantee sustained resistance. Certain neutralization techniques, particularly those framed in moral language, are systematically harder for models to resist. MoralRatBench will be released as an open-source evaluation tool.

When commenting on articles, please be friendly, welcoming, respectful and abide by the AIS eLibrary Discussion Thread Code of Conduct posted here.