提出新评估方法,让大模型更真实地模拟人类主观行为。
Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

- 用主观性系数区分任务难易,揭示准确率评估的局限性
- 设计自适应软标签训练,根据输入主观性调整标签范围
- 构建1.9万条数据的基准集,支持分布式行为评估
基于大语言模型的社会模拟是传统调研与行为实验的有力补充。当前主流评估方式依赖准确率,仅检查模型是否匹配人类单次观测结果,并以此硬标签训练模型。但人类行为具有内在主观性:同一情境下个体可能合理做出不同选择,单次观测仅为潜在响应分布的一个样本,导致准确率评估不可靠,硬标签训练误导模型。为此,本文提出主观性系数——一种基于熵的量化指标,用于区分客观任务(如编程)与主观任务(如社会模拟),并系统分析准确率评估与硬标签训练随主观性增强而失效的机制。基于此,提出主观性自适应软标签训练(SALT):将语义相近输入的观测结果聚合为软分布标签,聚类半径依据输入主观性动态调整;在近客观场景下,邻域缩小,SALT自然退化为标准单标签训练。此外,因现有数据集仅记录单一响应,无法支持分布评估,本文构建了包含19,300个情境、193名标注者和100个主观问题的基准集SUBJSIM。在真实设置下,仅用单次观测训练模型,却在完整响应分布上评估性能,验证了方法可行性。SUBJSIM上的实验表明,该方法显著优于传统范式。
原文摘要 · Abstract (English)
LLM-based social simulation is a promising complement to traditional methods such as surveys and behavioral experiments. A core question is how to evaluate the fidelity of LLM-simulated human behavior and optimize LLMs toward it. Prevailing practice evaluates by accuracy, checking whether the model selects the single response observed from a human, and trains the LLM to reproduce this hard label. However, human behavior is inherently subjective: the same person in the same situation may reasonably act differently, so an observed response is only one draw from an underlying response distribution, rendering accuracy-based evaluation unreliable and hard-label training misleading. To address these problems, we first introduce the subjectivity coefficient, an entropy-based quantity distinguishing objective tasks such as coding from subjective ones such as social simulation, and use it to systematically analyze how accuracy-based evaluation and hard-label training fail as subjectivity grows. Based on the subjectivity coefficient, we propose Subjectivity-Adaptive soft-Label Training (SALT): it pools observed outputs from semantically nearby inputs into soft distributional labels, with an aggregation radius adapted to the estimated subjectivity of each input; in the near-objective limit the neighborhood shrinks, so SALT naturally falls back to standard single-label training. Moreover, since existing datasets record only single observed responses and cannot support distributional evaluation, we construct SUBJSIM, a benchmark of 19,300 contexts covering 193 annotators and 100 subjective questions. Since real-world data typically provide only a single observation per input, our experiments train models from single observed outputs while evaluating them against the full response distributions, verifying feasibility in realistic settings. Results on SUBJSIM demonstrate the advantages of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。