arXiv:2605.04886cs.CL2026-05中稿 · LREC 2026

将社会科学数据纳入大模型评测,提升AI的泛化与社会相关性

BenCSSmark: Making the Social Sciences Count in LLM Research

  • 构建融合计算社会科学研究成果的评测基准BenCSSmark
  • 引入数十个严谨标注的社会科学数据集,增强模型泛化能力
  • 适合关注AI公平性、跨学科应用的研究者与政策制定者

本文主张,当前大语言模型(LLM)评测中社会科学任务的缺失,限制了模型评估与社会科学研究的进展。评测基准作为人工智能发展的重要工具,不仅衡量技术进步,更塑造研究方向与商业影响。尽管社会科学研究每年产出大量严谨标注、情境敏感的数据集,却极少被纳入主流评测框架。将这些数据融入评测设计,可显著提升AI模型的泛化性与鲁棒性。同时,基于社会科学任务训练的模型,有望在历史、社会学、政治学、经济学等多元领域表现更优。随着这些学科日益依赖大模型辅助研究,这一整合尤为紧迫。为此,本文提出BenCSSmark,一个由计算社会科学家标注的数据集构成的评测基准,旨在推动更稳健、透明且具社会意义的AI系统,并促进跨学科高效协作。

原文摘要 · Abstract (English)

This position paper argues that the under-representation of social science tasks in contemporary LLM benchmarks limits advances in both LLM evaluation and social scientific inquiry. Benchmarks -- standardized tools for assessing computational systems -- are pivotal in the development of artificial intelligence (AI), including large language models (LLMs). Benchmarks do more than measure progress -- they actively structure it, shaping reputations, research agendas, and commercial outcomes. Despite this central role, the social sciences are largely absent from mainstream evaluation frameworks, even though scholars in these fields generate dozens of rigorously annotated, context-sensitive datasets each year. Integrating this work into benchmark design could significantly improve the generalization and robustness of AI models. In turn, models trained on social scientific tasks would likely yield better performance on classic and contemporary tasks in disciplines as diverse as history, sociology, political science or economics. This is all the more pressing as these disciplines are quickly turning to LLMs for assistance. To address this gap, we introduce BenCSSmark, a benchmark composed of datasets annotated by computational social scientists. By integrating social scientific perspectives into benchmarking, BenCSSmark seeks to promote more robust, transparent, and socially relevant AI systems and to foster efficient collaboration.

大模型评测社会科学跨学科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。