arXiv:2606.26366cs.AIcs.CL2026-06

让大模型推理道德难题时更全面、更诚实,避免遗漏利益相关方和掩盖不确定性。

Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models

论文配图:Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models
图 1 · 摘自论文原文
  • 用五段式结构引导模型思考:主角、利益相关方、两步后果、不确定性、最终决定。
  • 使利益相关方遗漏率从31%降至1%以下,不确定性回避率从72%降到1-24%。
  • 无需训练,可直接提升多个大模型的伦理推理质量,适合需要可信决策的智能体应用。

标准链式思维在道德困境中存在两种缺陷:利益相关方坍缩(仅提及一个相关方)和不确定性压制(未明确未知或保留余地就做决定)。本文提出叙事性思维(NoT),通过系统提示将链式思维划分为五个部分:主角、利益相关方、两步后果、不确定性、最终承诺。NoT不需训练、参数或微调。在100个DailyDilemmas场景中,对来自三个供应商的四个生成器测试显示,该方法将利益相关方遗漏率从最高31%降至1%以下,不确定性压制率从最高72%降至1%-24%。控制实验排除了令牌消耗的影响;在相同预算下,NoT仍显著优于普通链式思维,提升幅度为+0.79至+0.90(利益相关方数量)和+0.65至+0.93(不确定性评分)。消融实验证明每部分指令均有独立贡献。以叙事性思维为起点进行文本梯度下降可进一步优化效果;跨厂商训练的裁判模型在所有指标上均优于同厂商模型。扩展至五轮多方辩论协议后,该框架将6%的僵局转化为95%共识,并在复现测试中实现100%联合收敛。生成的推理轨迹清晰呈现每个决策背后的参与者、后果与不确定性依据,为可靠智能体部署提供可审计基础。

原文摘要 · Abstract (English)

Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with a stake in the outcome) and uncertainty suppression (no explicit unknowns or hedges before committing to an action). We introduce narration-of-thought (NoT), a system prompt that structures chain-of-thought into five sections: protagonist, stakeholders, two-step consequences, uncertainty, then commitment. NoT adds no training, parameters, or fine-tuning. On 100 DailyDilemmas scenarios across four generators from three vendors, NoT cuts stakeholder collapse from up to 31% to under 1% and uncertainty suppression from up to 72% to 1-24% on every model. A matched-budget verbose-CoT control rules out token spend as the active ingredient; NoT retains Cliff's delta advantages of +0.79 to +0.90 on stakeholder count and +0.65 to +0.93 on uncertainty score for three of four generators, and a section ablation attributes each shift to its specific sub-instruction. Textual-gradient descent initialised at NoT improves the scaffold further; a cross-family training judge (different vendor from the generator) dominates an in-family one on every measured axis. Extended to a five-round multi-stakeholder debate protocol, the scaffold converts a 6% standoff into 95% full consensus on a calibration set and 100% combined convergence on a DailyDilemmas replication. The resulting traces externalise the stakeholders, consequences, and uncertainty grounding each commitment, providing an auditable substrate for dependable agentic deployment.

伦理推理链式思维大模型评估可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。