arXiv:2509.24841cs.CLcs.AI2025-09被引 1

提出分层纠错框架,提升通信研究自动化编码的可靠性与有效性。

A Hierarchical Error Framework for Reliable Automated Coding in Communication Research: Applications to Health and Political Communication

  • 将模型错误分为知识、推理、复杂度三层,逐层诊断并干预。
  • 平均准确率提升11.2个百分点,系统性误判显著减少。
  • 适用于健康与政治传播等场景,尤其适合中高难度任务。

自动化内容分析在传播学研究中日益重要,但将人工编码规模化为计算流程时,测量的可靠性和有效性面临挑战。本文提出分层误差校正(HEC)框架,将模型失败视为分层测量误差(知识盲区、推理局限、复杂度约束),并针对影响推断最严重的层级进行干预。该框架采用三阶段方法:跨层级系统性误差分析、匹配主导误差源的干预设计、基于统计检验的严格验证。在医疗专业分类(健康传播)和偏见检测(政治传播)及法律任务中评估,使用五种大型语言模型验证,结果显示平均准确率提升11.2个百分点(p < .001,McNemar检验),系统性误分类大幅降低。跨模型验证表明改进稳定(范围+6.8至+14.6个百分点),且在基线准确率50%-85%的任务中效果最佳。边界研究表明,在极高基线(>85%)或精度匹配任务中收益递减,明确了适用范围。通过将分层误差映射至效度威胁,提供透明、以测量为核心的诊断与报告蓝图,支持自动化编码在传播学及社会科学中的可靠应用。

原文摘要 · Abstract (English)

Automated content analysis increasingly supports communication research, yet scaling manual coding into computational pipelines raises concerns about measurement reliability and validity. We introduce a Hierarchical Error Correction (HEC) framework that treats model failures as layered measurement errors (knowledge gaps, reasoning limitations, and complexity constraints) and targets the layers that most affect inference. The framework implements a three-phase methodology: systematic error profiling across hierarchical layers, targeted intervention design matched to dominant error sources, and rigorous validation with statistical testing. Evaluating HEC across health communication (medical specialty classification) and political communication (bias detection), and legal tasks, we validate the approach with five diverse large language models. Results show average accuracy gains of 11.2 percentage points (p < .001, McNemar's test) and stable conclusions via reduced systematic misclassification. Cross-model validation demonstrates consistent improvements (range: +6.8 to +14.6pp), with effectiveness concentrated in moderate-to-high baseline tasks (50-85% accuracy). A boundary study reveals diminished returns in very high-baseline (>85%) or precision-matching tasks, establishing applicability limits. We map layered errors to threats to construct and criterion validity and provide a transparent, measurement-first blueprint for diagnosing error profiles, selecting targeted interventions, and reporting reliability/validity evidence alongside accuracy. This applies to automated coding across communication research and the broader social sciences.

自动化编码误差分析可靠性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。