arXiv:2609.05088cs.AIcs.CL2026-09被引 1

用对话分析法评估大模型道德推理的辩护质量,发现其论证优于事后解释。

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

论文配图:Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?
图 1 · 摘自论文原文
  • 设计四阶段对话协议,从推理过程和事后辩护两方面评估模型论证结构。
  • 9个前沿模型在200个高模糊性道德题中平均辩护得分远超最低标准,但理由充分性是短板。
  • 发现模型自相矛盾或前提错误等硬伤,适合关注AI伦理可解释性的研究者参考。

当前AI监管依赖真实答案进行验证,但何为合理AI行为本身存在争议,导致对大语言模型的道德推理及基于辩论的监督评估难以应对现实模糊性。本文提出一种可在模糊情境下运作的替代标准:通过基于沃尔顿论证方案理论与戈维埃论证严密性标准的四阶段辩证协议,衡量模型对其判断所作辩护的结构质量。该协议适应不同推理框架,突破多选题限制,同时考察判决前的推理与判决后的解释。在9个前沿模型与200个高模糊性道德选择题上,共生成6,778个经评委打分的评价单元,其二元失败判断的评委一致性达89.6%。所有模型在各维度上的辩护表现均显著高于基准线,失败集中于理由充分性与前提合理性,且与认知规避表达相关,而非论证长度。模型事前推理的辩护能力普遍优于事后解释,且各模型在解释中使用的论证方案与实际推理中采用的方案存在显著差异(每模型≥20%),尽管两者均以价值导向的实践推理为主。协议成功识别出严格不可辩护的缺陷(如自我矛盾、虚假前提),并揭示了撤回行为在AI对齐中的角色难以界定,提示需开展更贴近语境的评估。

原文摘要 · Abstract (English)

AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested. This leaves evaluation of moral reasoning in LLMs and debate-based oversight implicitly avoiding realistic ambiguity. We investigate an alternative standard designed to function despite such ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton's theory of argumentation schemes and Govier's criteria for argument cogency. The protocol is adaptive to different frames of reasoning, extends beyond multiple-choice framing, and treats both the reasoning that precedes a verdict and its post-hoc justification. Across nine frontier models and 200 high-ambiguity MoralChoice items -- $6,778$ judge-scored cells, validated against $89.6\%$ inter-judge agreement on the binary failure judgment -- models defend their reasoning well above the rubric minimum on every dimension. Failure mass concentrates on grounds and sufficiency, and correlates with epistemic hedging rather than argument length. Reasoning is better defended than post-hoc justification, on every model and every Govier dimension. The scheme a model presents in its justification differs from the one it reasoned with on a substantial share of dilemmas ($\geq 20\%$ per model), despite value-based practical reasoning dominating both tracks. The protocol catches strictly indefensible defences (self-contradiction, false premises), and it surfaces difficulties in characterizing the role of retraction in AI alignment, suggesting a need for more situated evaluations.

AI伦理论证分析大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。