arXiv:2606.12848cs.AIecon.GN2026-06被引 3

人类监督让AI社科研究更可靠,失败率从72%降至16%。

(Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable

论文配图:(Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable
图 1 · 摘自论文原文
  • 设计人机协作架构,限制AI只推理不执行数据任务。
  • 在四个数据集上实验,失败率从72%降到16%,显著降低。
  • 适合需要高可靠性、低偏差的社科研究者使用。

大型语言模型(LLMs)正被用于原本由专业研究人员完成的任务,如假设生成、模型设定选择和结论起草。我们提出,AI辅助研究的可靠性不仅取决于模型能力,更取决于人机间认知劳动的分配方式。通过基于预承诺、决策顺序、问责机制和注意力分配的「人机协同经济研究」(HLER)框架,在2×4因子实验中,280次完整研究运行显示,无约束多智能体基线的失败率达72%;而采用相同模型与代理分解、一致提示的HLER,仅16%失败。关键架构承诺包括:LLM仅推理不执行数据工作、数据与估计过程确定性处理、设置三道人工决策关卡。费舍尔精确检验表明失败率差异极显著(p<0.001)。在代表性最弱的清代人口登记数据上,可靠性提升最为明显,符合任务型生产模型与弗雷歇分布输出质量特征。80次消融实验显示,确定性计算与人工关卡独立贡献,且存在互补性。我们视HLER为研究支架而非自主科研智能体:显著减少失败,暴露残余弱点,防止不可靠结论进入发表环节。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for tasks once reserved for trained researchers, including hypothesis generation, specification choice, and drafting conclusions. We argue that the reliability of AI-assisted research depends not only on model capability, but also on how cognitive labour is structured between humans and machines. We study this problem through Human-in-the-Loop Economic Research (HLER), a decision architecture based on pre-commitment, decision sequencing, accountability, and attention allocation. In a pre-specified 2*4 factorial experiment with 280 complete research runs across four datasets, an unconstrained multi-agent baseline produced critical failures in 72% of runs. Using the same underlying model, the same agent decomposition, and identical prompts for the shared reasoning agents, HLER reduced the failure rate to 16% by imposing three architectural commitments: LLMs reason but do not execute data work, data and estimation are handled deterministically, and three human decision gates bind the workflow. Fisher's exact test rejects equality of failure rates at p<0.001. Reliability gains were largest on the least publicly represented dataset, a Qing-dynasty population register, consistent with a task-based production model with Frechet-distributed output quality. An 80-run ablation suggests that deterministic computation and human gates contribute independently, with exploratory evidence of complementarity. We interpret HLER as a research harness rather than an autonomous AI scientist: it sharply reduces failures, makes residual weaknesses more visible, and prevents unreliable claims from being advanced as publication-ready outputs.

人机协作社科研究可靠性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。