用贝叶斯方法评估大模型道德理解,发现其检出能力优于人类。
Beyond Human Judgment: A Bayesian Evaluation of LLMs' Moral Values Understanding
- 通过建模标注者分歧,区分人类固有分歧与模型认知不确定性。
- 在超10万文本上测试,大模型表现达人类前25%,误判率更低。
- 适合关注大模型伦理能力、可信评估的研究者参考。
大型语言模型如何理解道德维度?这是首个对主流语言模型进行大规模贝叶斯评估的研究。不同于以往使用确定性真值(多数或包含规则)的方法,本文建模标注者分歧,以捕捉固有的人类不确定性(aleatoric uncertainty)和模型领域敏感性(epistemic uncertainty)。我们评估了最佳语言模型(Claude Sonnet 4、DeepSeek-V3、Llama 4 Maverick),基于近700名标注者在超过10万篇文本(涵盖社交网络、新闻、论坛)上的25万+条标注数据。经过优化的GPU贝叶斯框架处理了超100万次模型查询,结果显示:大模型通常位于人类标注者前25%水平,整体准确率远超平均平衡准确率。更重要的是,模型产生的假阴性远少于人类,表明其具有更强的道德敏感性检测能力。
原文摘要 · Abstract (English)
How do Large Language Models understand moral dimensions compared to humans? This first large-scale Bayesian evaluation of market-leading language models provides the answer. In contrast to prior work using deterministic ground truth (majority or inclusion rules), we model annotator disagreements to capture both aleatoric uncertainty (inherent human disagreement) and epistemic uncertainty (model domain sensitivity). We evaluated the best language models (Claude Sonnet 4, DeepSeek-V3, Llama 4 Maverick) across 250K+ annotations from nearly 700 annotators in 100K+ texts spanning social networks, news and forums. Our GPU-optimized Bayesian framework processed 1M+ model queries, revealing that AI models typically rank among the top 25\% of human annotators, performing much better than average balanced accuracy. Importantly, we find that AI produces far fewer false negatives than humans, highlighting their more sensitive moral detection capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。