大模型编码政治事件时,仅靠准确率不够,需确保规则遵循性。
When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding

- 用定义+示例+规则增强提示,提升编码准确率
- 增强提示使宏观F1从0.457升至0.633
- 标签名或映射改变后,性能骤降,说明可靠性关键
高准确率并不意味着大模型是可靠的文本编码者。在社会科学研究中,常依赖专家编写的代码手册将文本转为结构化数据。本文研究政治事件编码任务,即模型需根据详细规则识别行为人对另一方的行动。比较仅使用标签名、简明定义及包含示例、事件模式指令和边界规则的增强引导。同时评估不同提示与检索方法。进一步测试在改变代码本顺序、标签名或标签-定义映射关系下的行为可靠性。结果显示,增强引导使均值根层级宏F1从0.457提升至0.633。具备定义信息的方法在移除有意义标签名后仍有效,但一旦标签-定义映射被重置,所有方法加权F1均未超过0.20。这表明需分别评估预测性能与规则遵循性。
原文摘要 · Abstract (English)
High accuracy does not necessarily make an LLM a faithful coder. This issue matters because many social-science studies rely on expert-written codebooks to turn text into structured data. We study political event coding, where a model must identify the action that one actor directs toward another under detailed coding rules. We compare label names alone with concise definitions and enriched guidance that adds examples, event-mode instructions, and boundary rules. We also evaluate alternative prompting and retrieval methods. We then test behavioral reliability under changes to codebook order, label names, and label-definition mappings. Enriched guidance raises mean root-level macro-F1 from 0.457 to 0.633. Methods with access to definitions remain effective when meaningful label names are removed, but no evaluated method exceeds 0.20 weighted F1 after the label-definition mapping is reassigned. These results motivate separate evaluation of predictive performance and adherence to the supplied coding rules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。