arXiv:2505.08106cs.CLcs.AI2025-05被引 1

测试大模型能否像人一样分析伦理困境,发现它们结构上胜过普通人但缺乏深层理解。

Are LLMs complicated ethical dilemma analyzers?

  • 构建196个真实伦理困境数据集,分五部分对比专家与模型输出。
  • GPT-4o-mini在各环节表现最稳定,但所有模型难处理历史背景和复杂对策。
  • 适合研究伦理推理、AI可解释性或人机决策的学者和开发者。

大型语言模型(LLMs)是否能模拟人类的道德推理并作为人类判断的可信代理,是一个开放问题。为此,我们构建了一个包含196个真实世界伦理困境及其专家意见的基准数据集,每个案例分为五个结构化部分:引言、关键因素、历史理论视角、解决策略和核心启示。同时收集非专家人类回答用于对比,仅限于‘关键因素’部分。我们评估了多个前沿模型(GPT-4o-mini、Claude-3.5-Sonnet、Deepseek-V3、Gemini-1.5-Flash),采用基于BLEU、Damerau-Levenshtein距离、TF-IDF余弦相似度和通用句子编码器相似度的复合指标框架,权重通过反向排序对齐与成对AHP分析计算,实现对模型输出与专家回答的细粒度比较。结果显示,大模型在词汇与结构匹配度上普遍优于非专家人类,其中GPT-4o-mini在各部分表现最一致。然而,所有模型在历史背景把握和提出精细解决方案方面均存在困难,这些任务需要上下文抽象能力。人类回答虽结构松散,但有时在语义相似度上可达相当水平,表明其具有直觉性道德推理能力。结果揭示了大模型在伦理决策中的优势与当前局限。

原文摘要 · Abstract (English)

One open question in the study of Large Language Models (LLMs) is whether they can emulate human ethical reasoning and act as believable proxies for human judgment. To investigate this, we introduce a benchmark dataset comprising 196 real-world ethical dilemmas and expert opinions, each segmented into five structured components: Introduction, Key Factors, Historical Theoretical Perspectives, Resolution Strategies, and Key Takeaways. We also collect non-expert human responses for comparison, limited to the Key Factors section due to their brevity. We evaluate multiple frontier LLMs (GPT-4o-mini, Claude-3.5-Sonnet, Deepseek-V3, Gemini-1.5-Flash) using a composite metric framework based on BLEU, Damerau-Levenshtein distance, TF-IDF cosine similarity, and Universal Sentence Encoder similarity. Metric weights are computed through an inversion-based ranking alignment and pairwise AHP analysis, enabling fine-grained comparison of model outputs to expert responses. Our results show that LLMs generally outperform non-expert humans in lexical and structural alignment, with GPT-4o-mini performing most consistently across all sections. However, all models struggle with historical grounding and proposing nuanced resolution strategies, which require contextual abstraction. Human responses, while less structured, occasionally achieve comparable semantic similarity, suggesting intuitive moral reasoning. These findings highlight both the strengths and current limitations of LLMs in ethical decision-making.

伦理推理大模型评估人机对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。