arXiv:2602.13274cs.AIcs.CL2026-02被引 4

对比11种提示策略,发现简单示例提示更安全高效

ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs

  • 用11种提示范式测试4类大模型的道德与安全表现
  • 示例引导的简洁提示比复杂推理更稳定且节省计算资源
  • 提出统一评分标准,适合安全对齐研究者使用

提示设计显著影响大语言模型的道德能力与安全对齐,但现有评估分散于不同数据集和模型。我们提出ProMoral-Bench,一个统一基准,评估11种提示范式在四个大模型家族中的表现。基于ETHICS、Scruples、WildJailbreak及新提出的鲁棒性测试ETHICS-Contrast,采用自研的统一道德安全得分(UMSS)进行评估。结果表明,紧凑的示例引导提示优于复杂的多阶段推理,具有更高的UMSS得分和更强鲁棒性,且消耗更少令牌。多轮推理在扰动下易失效,而少样本示例能持续提升道德稳定性与越狱抵抗能力。ProMoral-Bench建立了标准化的、成本可控的提示工程框架。

原文摘要 · Abstract (English)

Prompt design significantly impacts the moral competence and safety alignment of large language models (LLMs), yet empirical comparisons remain fragmented across datasets and models.We introduce ProMoral-Bench, a unified benchmark evaluating 11 prompting paradigms across four LLM families. Using ETHICS, Scruples, WildJailbreak, and our new robustness test, ETHICS-Contrast, we measure performance via our proposed Unified Moral Safety Score (UMSS), a metric balancing accuracy and safety. Our results show that compact, exemplar-guided scaffolds outperform complex multi-stage reasoning, providing higher UMSS scores and greater robustness at a lower token cost. While multi-turn reasoning proves fragile under perturbations, few-shot exemplars consistently enhance moral stability and jailbreak resistance. ProMoral-Bench establishes a standardized framework for principled, cost-effective prompt engineering.

提示工程模型安全道德推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。