量化大模型观点多样性,提出可自动评估的基准
Benchmarking Overton Pluralism in LLMs
- 用集合覆盖度定义观点多元性,构建可衡量指标
- 模型平均得分0.35-0.41,最优为DeepSeek V3,远低于理论上限1.0
- 自动化基准与人工判断高度一致(ρ=0.88),适合高效评估
我们提出OVERTONBENCH,一个用于测量大语言模型中观点多元性的新框架——即模型输出中多样观点的呈现程度。我们(一)将观点多元性形式化为集合覆盖指标(OVERTONSCORE),(二)开展大规模美国代表性人群研究(N=1208,60个问题,8个LLM),(三)开发一个能精准复现人工判断的自动化基准。结果显示,模型平均OVERTONSCORE为0.35–0.41,其中DeepSeek V3表现最佳,但均显著低于理论最大值1.0,表明仍有巨大改进空间。由于大规模人工评估成本高、耗时长,可扩展的评估工具至关重要。因此,我们提出的自动化基准与人工判断具有高度相关性(ρ=0.88),可在不替代人类评估的前提下提供实用代理。本工作将多元对齐从规范目标转变为可衡量基准,为构建更具多元性的大模型奠定系统化基础。
原文摘要 · Abstract (English)
We introduce OVERTONBENCH, a novel framework for measuring Overton pluralism in LLMs--the extent to which diverse viewpoints are represented in model outputs. We (i) formalize Overton pluralism as a set coverage metric (OVERTONSCORE), (ii) conduct a large-scale U.S.-representative human study (N = 1208; 60 questions; 8 LLMs), and (iii) develop an automated benchmark that closely reproduces human judgments. On average, models achieve OVERTONSCOREs of 0.35--0.41, with DeepSeek V3 performing best; yet all models remain far below the theoretical maximum of 1.0, revealing substantial headroom for improvement. Because repeated large-scale human studies are costly and slow, scalable evaluation tools are essential for model development. Hence, we propose an automated benchmark that achieves high rank correlation with human judgments ($ρ= 0.88$), providing a practical proxy without replacing human assessment. By turning pluralistic alignment from a normative aim into a measurable benchmark, our work establishes a foundation for systematic progress toward more pluralistic LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。