arXiv:2606.28574cs.CLcs.AI2026-06被引 1

提出粒度校准法,让大模型编码更符合理论构念本质。

Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs

  • 将构念分解为句级成分,逐项用文本证据验证
  • 通过理论导出的显式规则组合结果,避免黑箱决策
  • 可识别错误是遗漏成分还是误判邻近构念,适合理论验证

当大型语言模型(LLM)像人工标注者一样对文本中的构念进行编码时,这种一致性使该模型成为可靠的编码工具。然而,可靠性并未解决构念效度问题。该工具可能缺乏理论意识,仅通过满足某些外在相关性而得出编码结果,却未满足理论所要求的任何条件,而当前方法无法区分这种虚假测量与真实测量。本文提出粒度校准(grain calibration)方法,以填补这一空白。该方法将构念分解为句级组件,针对每项组件在文本中提取证据进行测试,并通过一个显式的、基于理论推导的规则整合结果。由于规则是明确陈述而非隐含于单一处理流程中,其结构本身即成为过程的证据,而非输出结果。该方法可揭示哪些组件决定了最终编码,当编码错误时,也能判断是遗漏了某个成分,还是将相邻构念误认为目标构念。验证重点从比较模型输出与人工标注,转变为证明模型确实在执行理论所规定的构念过程。

原文摘要 · Abstract (English)

When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder. Yet reliability leaves construct validity untouched. The instrument may be theory-naive, reaching the code through a correlate that meets none of the demands the construct's theory makes, and no current method tells that apart from genuine measurement. We propose grain calibration as a method that closes the gap. It decomposes a construct into clause-level components, tests each against the text with extractive evidence, and combines the results through an explicit, theory-derived rule. Because the rule is stated rather than lodged in one opaque pass, its structure is evidence about the process rather than the output. It shows which components settled a code, and, when the code is wrong, whether a component was missed or an adjacent construct mistaken for it. Validation shifts from scoring an instrument's outputs against an annotator to showing that the instrument runs on the construct its theory specifies.

大模型评估构念效度理论验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。