用寓言故事测试大模型的抽象道德推理能力,发现其易受干扰且常自相矛盾。
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
- 基于历史寓言构建多选题,通过精心设计干扰项考察深层道德推断。
- 顶尖模型在不同表述下自相矛盾率达20%,暴露其依赖表面模式而非真实推理。
- 增强推理能力未能提升表现,说明规模仍是关键,而非逻辑训练。
随着大模型在标准阅读理解任务中表现优异,评估其复杂抽象推理与推断能力成为新焦点。基于文学的故事类基准因其丰富的叙事与道德内涵,为评估深层理解提供了有力框架。本文提出 MORABLES,一个由人类验证的基准,源自历史文献中的寓言与短篇故事。主要任务为针对道德推断设计的多选题,干扰项经精心构造,迫使模型超越浅层抽取式回答。为进一步检验模型鲁棒性,引入对抗性变体以揭示因数据污染等问题导致的漏洞与捷径。结果表明,虽大模型优于小模型,但仍易受对抗干扰,常依赖表面模式而非真正道德推理。这种脆弱性导致显著自相矛盾:最佳模型在不同道德表述下会否定自身答案,比例达约20%。有趣的是,增强推理能力的模型仍未缩小差距,暗示性能提升主要源于规模,而非推理能力本身。
原文摘要 · Abstract (English)
As LLMs excel on standard reading comprehension benchmarks, attention is shifting toward evaluating their capacity for complex abstract reasoning and inference. Literature-based benchmarks, with their rich narrative and moral depth, provide a compelling framework for evaluating such deeper comprehension skills. Here, we present MORABLES, a human-verified benchmark built from fables and short stories drawn from historical literature. The main task is structured as multiple-choice questions targeting moral inference, with carefully crafted distractors that challenge models to go beyond shallow, extractive question answering. To further stress-test model robustness, we introduce adversarial variants designed to surface LLM vulnerabilities and shortcuts due to issues such as data contamination. Our findings show that, while larger models outperform smaller ones, they remain susceptible to adversarial manipulation and often rely on superficial patterns rather than true moral reasoning. This brittleness results in significant self-contradiction, with the best models refuting their own answers in roughly 20% of cases depending on the framing of the moral choice. Interestingly, reasoning-enhanced models fail to bridge this gap, suggesting that scale - not reasoning ability - is the primary driver of performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。