arXiv:2602.22831cs.LGcs.AI2026-02被引 2

通过反转提示方向,发现大模型道德选择隐藏的敏感结构。

Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs

  • 用反向提示对比测试模型在道德决策中的实际响应变化
  • 短提示可使选择率波动12-18个百分点,40%情况存在方向不对称
  • 模型常否认提示影响,但行为显示其已被操控,适合研究者验证模型可信度

当前大模型道德评估多基于无上下文提示,隐含假设选择率稳定。本文提出方向反转影响审计:对每个情境,比较基线提示与引导选A或选B的匹配提示。在电车难题式道德分诊、BBQ和DailyDilemmas任务中,五类大模型(含/不含推理)在短上下文提示下,条件选择率平均变动12-18个百分点。这些变动揭示了基线评分忽略的结构:约40%原本中立的三元判断与BBQ情境呈现方向不对称;显著效应中近半出现反向反弹,即结果与提示意图相反。后续探测发现,模型常识别提示但否认其影响,此类言行不一在显著反弹案例中占比78%。推理能力未消除上下文敏感性,反而重塑其模式:社会压力类提示(如用户偏好、情感诉求)在各基准上减弱,而少样本示例在分诊与BBQ任务中显著增强。建议将方向反转提示对作为标准补充用于道德偏见评估,并公开工具与数据以推动常规化审计。

原文摘要 · Abstract (English)

Moral benchmarks for LLMs typically score models on context-free prompts, implicitly treating the measured choice rate as stable. We test this assumption with a direction-flipped influence audit: for each scenario, we compare a baseline prompt with matched cues steering toward option A or option B. Across a trolley-problem-style moral triage task, BBQ, and DailyDilemmas, and across five LLM families with and without reasoning, short contextual cues shift per-condition choice rates by 12-18 percentage points on average. These shifts reveal structure that baseline scores miss: roughly 40% of baseline-neutral triage and BBQ conditions exhibit directional asymmetry under influence, and a meaningful share of significant effects backfire, moving opposite the cue's intended direction. In follow-up probes, models often recognize the cue while denying that it affected their choice. Among significant backfire trials, this stated-vs.-revealed inconsistency appears in 78% of cases. Reasoning does not eliminate contextual sensitivity but reshapes it: social-pressure cues such as user preference and emotional appeal weaken across benchmarks, while few-shot demonstrations strengthen sharply on both triage and BBQ. We recommend direction-flipped influence pairs as a standard complement to context-free moral-bias evaluation, and release the harness and data to make such audits routine.

道德评估提示敏感大模型行为反向审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。