提出RADIUS评估框架,系统衡量大模型生成问卷回答的排序与分布一致性。
RADIUS: Ranking, Distribution, and Significance - A Comprehensive Alignment Suite for Survey Simulation
- 构建二维评估体系:同时考察选项排序和分布匹配度。
- 引入显著性检验,确保评估结果可统计验证。
- 适合需要真实用户偏好模拟的决策类应用研究者使用。
利用大语言模型(LLM)进行调查模拟正成为大规模生成类人响应的强大工具。以往研究采用来自其他领域的评价指标,这些指标往往随意、零散且非标准化,导致结果难以比较。此外,现有指标主要关注准确率或分布匹配,忽略了排名对齐这一关键维度。实践中,即使模拟结果准确率高,也可能未能捕捉到人类最偏好的选项——这对决策类应用至关重要。我们提出RADIUS,一个全面的二维对齐评估套件,涵盖:1)排名对齐;2)分布对齐,并均配备统计显著性检验。RADIUS揭示了现有指标的局限性,使调查模拟评估更有效,提供开源实现以支持可复现、可比的评测。
原文摘要 · Abstract (English)
Simulation of surveys using LLMs is emerging as a powerful application for generating human-like responses at scale. Prior work evaluates survey simulation using metrics borrowed from other domains, which are often ad hoc, fragmented, and non-standardized, leading to results that are difficult to compare. Moreover, existing metrics focus mainly on accuracy or distributional measures, overlooking the critical dimension of ranking alignment. In practice, a simulation can achieve high accuracy while still failing to capture the option most preferred by humans - a distinction that is critical in decision-making applications. We introduce RADIUS, a comprehensive two-dimensional alignment suite for survey simulation that captures: 1) RAnking alignment and 2) DIstribUtion alignment, each complemented by statistical Significance testing. RADIUS highlights the limitations of existing metrics, enables more meaningful evaluation of survey simulation, and provides an open-source implementation for reproducible and comparable assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。