用大模型当多个评审之一,提升政治立场测量在数据稀疏地区的效果。
The Model as One Rater Among Several: Measuring Political Positions in Data-Sparse Regions with a Language-Model Panel
- 将大语言模型视为一个有误差的评审,通过多人打分聚合提升可靠性。
- 加入轴定义后评分平均变动1.8点,评审间一致性显著提高。
- 适用于西方工具失效的地区,尤其适合中东北非等数据稀缺区域。
现有政治立场测量工具多基于西方政党体系构建,在非西方情境下表现不佳甚至无效。本文提出一种新方法:将大语言模型视为一个可错的评审员,纳入由多个评审构成的面板中,类似专家调查中的多人评估。该方法包含适用性规则(区分零分与空白)、评分面板设计及分离言论与行为的透镜系统。结果显示,在固定定义条件下,引入书面轴定义使评分均值变动1.8点(21点量表),评审间绝对差距从2.81降至2.50,相关系数由0.81升至0.89;跨八家实验室九个模型,克里彭多夫阿尔法达0.86,且随评审人数从五人增至九人保持稳定。虽无法验证正确性,但分歧本身具信息价值——最严重分歧指向语义参照问题,三分之二可通过解释差异而非错误解决。研究公开全部工具与数据,案例聚焦中东与北非,方法可推广至其他标准工具覆盖不足的地区。
原文摘要 · Abstract (English)
Most tools for measuring political positions, manifesto coding, expert surveys, text-scaling models, were built and validated on Western party systems, and outside that setting they work poorly, and often not at all. This paper is an attempt at a method for those settings. It treats a large language model not as a measurement device but as a single, fallible rater in a panel, roughly the way an expert survey treats one expert: the value comes from pooling many judges rather than trusting any one of them. I describe the panel, an applicability rule that keeps a score of zero distinct from a blank, and a lens system that separates what an actor says from what it does. I report three results. First, holding a definition-free round fixed, adding written axis definitions moves scores by a mean of 1.8 points on a 21-point scale and tightens agreement between raters (mean absolute gap 2.81 to 2.50; r 0.81 to 0.89); they make two independent raters agree more closely, which an arbitrary steer would not. Second, across nine models from eight laboratories in two countries, Krippendorff's alpha is 0.86 on both an interval and an ordinal metric, and it stayed put as the panel grew from five raters to nine. That is reliability, the reproducibility of a reading, and not validity, its correctness. Third, where the panel does disagree, the disagreement is informative: the sharpest split, a full-scale divergence on an actor's stance toward its state's foundational order, points to a referent problem, and a blind triple-coding puts about two-thirds of it down to interpretation rather than error. I try to be plain about what the method can't do, including the human validation it still lacks, and I release the instrument and data in full. The worked example is the Middle East and North Africa, but I'd expect the method to carry to any region these standard tools leave out.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。