arXiv:2412.16783cs.CL2024-12被引 2

构建统一数据标准,评估大模型在政治与人口视角上的对齐程度

SubData: Bridging Heterogeneous Datasets to Enable Theory-Driven Evaluation of Political and Demographic Perspectives in LLMs

  • 提出SubData工具库,统一异构数据集格式
  • 基于理论框架测试不同立场模型对特定群体内容的分类差异
  • 开源可扩展,适合研究大模型偏见与社会影响的学者

随着大语言模型能力不断提升,研究人员开始探索其在主观任务中的潜力。尽管已有研究证明大模型可对齐多样人类观点,但在下游任务(如仇恨言论检测)中评估这种对齐仍因研究间数据集不一致而困难。为此,本文提出一个两步框架:(1) 引入SubData——一个开源Python库,用于标准化异构数据集以评估大模型的观点对齐;(2) 提出一种基于理论的方法,利用该库测试不同立场对齐的大模型(如不同政治倾向)如何分类针对特定人口群体的内容。SubData的灵活映射与分类体系支持多样化研究需求,区别于现有资源。我们通过示例应用展示其使用方式,并邀请贡献者将初始版本扩展为多维度基准套件,用于评估大模型在自然语言处理任务中对观点对齐的表现。

原文摘要 · Abstract (English)

As increasingly capable large language models (LLMs) emerge, researchers have begun exploring their potential for subjective tasks. While recent work demonstrates that LLMs can be aligned with diverse human perspectives, evaluating this alignment on downstream tasks (e.g., hate speech detection) remains challenging due to the use of inconsistent datasets across studies. To address this issue, in this resource paper we propose a two-step framework: we (1) introduce SubData, an open-source Python library designed for standardizing heterogeneous datasets to evaluate LLMs perspective alignment; and (2) present a theory-driven approach leveraging this library to test how differently-aligned LLMs (e.g., aligned with different political viewpoints) classify content targeting specific demographics. SubData's flexible mapping and taxonomy enable customization for diverse research needs, distinguishing it from existing resources. We illustrate its usage with an example application and invite contributions to extend our initial release into a multi-construct benchmark suite for evaluating LLMs perspective alignment on natural language processing tasks.

大模型对齐政治偏见数据标准化社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。