用大模型全流程模拟社会调查,评估其对人类偏好的对齐效果。
AlignSurvey: A Comprehensive Benchmark for Human Preferences Alignment in Social Surveys
- 构建全链条社会调查模拟框架,涵盖角色建模到回答生成。
- 基于4.4万条跨国家访谈与40万份问卷,覆盖多元人群。
- 提供公平性、一致性等多维度评估指标,适合政策研究者使用。
通过社会调查理解人类态度、偏好和行为对学术研究与政策制定至关重要。然而传统调查面临固定题型、成本高、适应性差及跨文化等价难等问题。尽管近期研究尝试用大语言模型(LLMs)模拟调查回应,但多数局限于结构化问题,忽略完整调查流程,且因训练数据偏差可能低估边缘群体。我们提出AlignSurvey,首个系统性复现并评估大模型在完整社会调查流程中表现的基准。它定义了四个与调查关键阶段对应的任务:社会角色建模、半结构化访谈建模、态度立场建模与调查回应建模,并提供任务特定评估指标,衡量个体与群体层面的对齐精度、一致性与公平性,尤其关注人口多样性。为支持此基准,我们构建了多层级数据架构:(i) 社会基础语料库,含44,000+跨国家访谈对话与400,000+结构化调查记录;(ii) 全流程调查数据集,包括专家标注的AlignSurvey-Expert (ASE) 及两个具有全国代表性的调查数据集,用于跨文化评估。我们发布SurveyLM系列模型,通过对开源大模型进行两阶段微调获得,并提供参考模型以评估领域对齐性能。所有数据集、模型与工具均开源至GitHub与HuggingFace,助力透明、负责任的研究。
原文摘要 · Abstract (English)
Understanding human attitudes, preferences, and behaviors through social surveys is essential for academic research and policymaking. Yet traditional surveys face persistent challenges, including fixed-question formats, high costs, limited adaptability, and difficulties ensuring cross-cultural equivalence. While recent studies explore large language models (LLMs) to simulate survey responses, most are limited to structured questions, overlook the entire survey process, and risks under-representing marginalized groups due to training data biases. We introduce AlignSurvey, the first benchmark that systematically replicates and evaluates the full social survey pipeline using LLMs. It defines four tasks aligned with key survey stages: social role modeling, semi-structured interview modeling, attitude stance modeling and survey response modeling. It also provides task-specific evaluation metrics to assess alignment fidelity, consistency, and fairness at both individual and group levels, with a focus on demographic diversity. To support AlignSurvey, we construct a multi-tiered dataset architecture: (i) the Social Foundation Corpus, a cross-national resource with 44K+ interview dialogues and 400K+ structured survey records; and (ii) a suite of Entire-Pipeline Survey Datasets, including the expert-annotated AlignSurvey-Expert (ASE) and two nationally representative surveys for cross-cultural evaluation. We release the SurveyLM family, obtained through two-stage fine-tuning of open-source LLMs, and offer reference models for evaluating domain-specific alignment. All datasets, models, and tools are available at github and huggingface to support transparent and socially responsible research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。