arXiv:2608.27049cs.CLcs.CY2026-08中稿 · EMNLP

用自动化方法检测社科论文研究设计并评估质量,提升政策决策可信度。

Research Design Tracking and Assessment for the Social Sciences

  • 构建基于多轮RAG的对话系统,自动识别六类反事实研究设计
  • 发现段落长度解释52%-66%性能差异,是关键影响因素
  • 机器难易与专家分歧不一致,揭示任务存在双重难度来源

社会科学中因果研究设计的可靠评估对证据驱动的政策制定至关重要,但目前完全依赖人工专家分析。本文提出自动化研究设计追踪与评估(ARDTrA),旨在检测论文使用的研究设计并评估其应用质量。我们构建了一个由专家标注的数据集,涵盖六类反事实研究设计,并采用基于多轮RAG的对话管道进行评估。在四种检索策略、四种大模型和六种嵌入模型下,结果表明段落长度是性能的主要驱动因素,解释了52%-66%的方差。进一步按研究设计分析发现,人类与机器的难度并不一致:系统最难处理的设计并非专家分歧最大的,暗示任务存在两个独立的难度来源。

原文摘要 · Abstract (English)

Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on manual expert analysis. We introduce Automated Research Design Tracking and Assessment (ARDTrA), a task that involves detecting the research design used in a paper and assessing the quality of its application. We create an expert-annotated dataset of papers covering six families of counterfactual research designs and evaluate the task using a multi-turn RAG-based conversational pipeline. Across four retrieval strategies, four LLMs and six embedding models, we find that passage length is the main driver of performance, explaining 52-66% of the variance. A per-research-design analysis also shows that human and machine difficulty do not align: the designs that prove hardest for the system are not those on which expert annotators disagree most, pointing to two independent sources of task difficulty.

研究设计自动化评估社会科学LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。