arXiv:2606.15566cs.CLcs.AI2026-06

用大模型检测科学论文中对贝叶斯模型的立场,准确率达76%以上。

LLM-Assisted Stance Detection in Scientific Discourse: A Test Case in Bayesian Cognitive Science

论文配图:LLM-Assisted Stance Detection in Scientific Discourse: A Test Case in Bayesian Cognitive Science
图 1 · 摘自论文原文
  • 基于理论构建编码手册,通过诊断筛选优化零样本提示
  • 在6858条引文上实现0.76的联合可靠性,文章级排名稳定度达0.96
  • 发现低阶感知运动研究更倾向现实主义,验证了学界长期直觉

定性编码是社会科学核心,但专家标注难以扩展。大语言模型(LLM)可提供补充,但在解释性强、理论负载重且间接表达的构念中需谨慎验证。本研究聚焦于判断作者是否将贝叶斯模型视为心理与神经机制的描述(现实主义)或仅是数学工具(工具主义)。方法结合理论驱动编码手册、专家标注参考数据、诊断门控提示优化搜索,生成适用于GPT-5.1、Claude Sonnet 4.6和Gemini 3 Pro Preview三款前沿模型的共享零样本提示,并进行多评者信度分析。最终提示在保留测试集上获得0.76的联合可靠性(谐均值ICC=0.79,α=0.74),所有诊断指标达标。部署于210篇论文中的6,858条引文,三模型在引文层面达成显著一致性(ICC=0.80;α=0.76;联合=0.78),文章级排名稳定性极高(相关系数r=0.96–0.97)。语料整体偏弱现实主义,但多数文章立场不单一:仅1.4%为单一立场,59.5%跨越四个及以上等级。低阶感知/运动类文章现实主义得分比高阶认知类高出8.8分(p<0.001,d=0.60),量化了长期存在的质性直觉。本研究以专家主导案例呈现,框架旨在推广至类似理论复杂任务,而非所有定性分析。

原文摘要 · Abstract (English)

Qualitative coding is central to social science, but expert annotation is difficult to scale. LLMs offer a possible extension, yet require careful validation when the target construct is interpretive, theoretically loaded, and only indirectly expressed. We study this problem in a difficult case: detecting whether authors treat Bayesian models as descriptions of mental and neural mechanisms (realism) or as useful mathematical tools (instrumentalism). Our method combines a theory-driven codebook, expert-coded reference annotations, a diagnostic-gated prompt-optimization search yielding a shared zero-shot prompt for three frontier LLMs (GPT-5.1, Claude Sonnet 4.6, Gemini 3 Pro Preview), and multi-rater reliability analysis. The final prompt achieved a held-out combined reliability score of 0.76 (harmonic mean of ICC = 0.79 and $α$ = 0.74), with all diagnostics satisfied. Deployed on 6,858 quotes from 210 articles, the three LLMs reached substantial quote-level agreement (ICC = 0.80; $α$ = 0.76; combined = 0.78) and near-perfect article-level rank stability ($r$ = 0.96-0.97 across rater pairs). The corpus was predominantly weakly realist, but article-level stances were rarely uniform: only 1.4% of articles used a single band, while 59.5% spanned four or more. Low-level perception/motor articles scored 8.8 Realism points higher than high-level cognition articles ($p < .001$, $d = 0.60$), quantifying a long-held qualitative intuition. We present this as an expert-led case study; the framework is intended to generalize to similar theoretically demanding tasks, not to all qualitative analysis.

立场检测大模型应用贝叶斯认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。