用30个真实伦理场景测试16个大模型,发现其伦理推理准确率达86.7%。
Advancing Automated Ethical Profiling in SE: a Zero-Shot Evaluation of LLM Reasoning
- 零样本评估16个大模型在30个伦理场景中的推理能力
- 平均理论一致性率达73.3%,道德可接受性判断准确率86.7%
- 适合关注AI伦理集成的软件工程开发者和研究者
大语言模型(LLMs)正被广泛应用于软件工程工具中,处理超出代码生成的任务,如不确定性下的判断与重要伦理情境中的推理。本文提出一个完全自动化的框架,在零样本设置下评估16个大模型在30个真实世界伦理场景中的伦理推理能力。每个模型需识别最适用的伦理理论、评估行为的道德可接受性,并解释理由。结果通过模型间一致性指标与专家伦理学家判断对比。结果显示,大模型平均理论一致性率(TCR)为73.3%,道德可接受性二值一致率(BAR)达86.7%,差异主要集中在伦理模糊案例中。对自由文本解释的定性分析显示,尽管词汇表达多样,但概念上存在强一致性。这些发现支持大模型作为软件工程流水线中伦理推理引擎的可行性,实现可扩展、可审计、自适应的用户对齐伦理推理。本研究聚焦于更广义伦理画像流程中的‘伦理解释器’组件,验证当前大模型是否具备足够的解释稳定性与理论一致性推理能力以支撑自动化画像。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly integrated into software engineering (SE) tools for tasks that extend beyond code synthesis, including judgment under uncertainty and reasoning in ethically significant contexts. We present a fully automated framework for assessing ethical reasoning capabilities across 16 LLMs in a zero-shot setting, using 30 real-world ethically charged scenarios. Each model is prompted to identify the most applicable ethical theory to an action, assess its moral acceptability, and explain the reasoning behind their choice. Responses are compared against expert ethicists' choices using inter-model agreement metrics. Our results show that LLMs achieve an average Theory Consistency Rate (TCR) of 73.3% and Binary Agreement Rate (BAR) on moral acceptability of 86.7%, with interpretable divergences concentrated in ethically ambiguous cases. A qualitative analysis of free-text explanations reveals strong conceptual convergence across models despite surface-level lexical diversity. These findings support the potential viability of LLMs as ethical inference engines within SE pipelines, enabling scalable, auditable, and adaptive integration of user-aligned ethical reasoning. Our focus is the Ethical Interpreter component of a broader profiling pipeline: we evaluate whether current LLMs exhibit sufficient interpretive stability and theory-consistent reasoning to support automated profiling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。