arXiv:2607.19834cs.CL2026-07

构建日常价值困境测评集,评估大模型的伦理对齐能力

D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios

论文配图:D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios
图 1 · 摘自论文原文
  • 用人类与大模型协作生成1万条真实生活价值冲突场景
  • 8个主流大模型在多维度价值对齐上表现差异显著
  • 适合研究模型伦理、价值观对齐的学者使用

随着大语言模型在现实场景中的广泛应用,其输出所隐含的价值观至关重要。然而,现有评估基准在日常情境下的价值冲突覆盖不足,且评价形式过于简单,难以有效评估模型的价值对齐。为此,我们提出D2VBench,一个包含10,000个真实日常困境场景的价值对齐评测基准,通过人类与大模型的多阶段协作构建,基于158个手工标注的细粒度价值概念。针对该基准,我们设计了一种融合多项选择与开放问答的混合评估范式,并对8个主流大语言模型进行了全面评估。实验结果表明,D2VBench具有高可靠性与鲁棒性,能有效反映模型在不同价值类别与维度上的对齐水平,为价值对齐研究提供了更真实、更精细的工具。数据集已开源:https://github.com/tjunlp-lab/D2VBench。

原文摘要 · Abstract (English)

With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in daily scenarios involving multiple value conflicts and simplistic evaluation formalisms that fail to assess LLMs' value alignment. To address these issues, we propose D2VBench, a value alignment benchmark comprising 10,000 instances of real daily dilemma scenarios constructed through a multi-stage collaboration between LLMs and humans, grounded in 158 manually annotated fine-grained value concepts. For evaluation on the benchmark, we present a hybrid evaluation paradigm that integrates multiple-choice questions with open-ended questions. We conduct comprehensive evaluations on eight mainstream LLMs. Experimental results demonstrate that D2VBench exhibits high reliability and robustness, effectively reflecting the LLMs' alignment across different value categories and dimensions, and providing a more realistic and fine-grained tool for research on value alignment. The dataset is available at https://github.com/tjunlp-lab/D2VBench.

价值对齐大模型评估伦理测试基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。