arXiv:2607.08731cs.CLcs.AI2026-07

用可复现的审计法检验国产大模型能否真实测量社会价值观。

Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA

  • 通过理论重构代码本并量化还原度,评估模型是否真懂概念。
  • 90亿参数的葡萄牙模型在权威性编码上表现接近八到十三倍大的开源模型。
  • 发现模型只解释了理论预期一半的表现,说明需校准才能当科学仪器用。

国家语言模型正成为公共资助的知识基础设施。公共所有权、语言专精和开放权重带来可信假定。这类由语言社区共建共用的模型,看似是衡量该群体言论与价值的天然工具。但其有效性在发布时未被验证。现有对大模型的评估多为任务特定,仅止于与人工标注者的一致性。一致性无法区分模型是真正测量某一构念,还是仅通过表面相关物达成匹配。本文在有利案例中审计这一假定:葡萄牙90亿参数的公开模型AMALIA,用于分析欧洲葡萄牙语中的权威性道德基础。我们提出“还原差距”(recovery gap)作为审计方法:将代码本按理论定义拆分为条款,依理论规则重组,再测量原始提示下理论能复现多少编码表现。在一项预注册的跨语言(英译欧葡)样本研究中,AMALIA与受训标注者的吻合度比八至十三倍大的开放模型仅低六分。然而,还原差距显示,仅约一半的权威性编码表现可归因于理论。更大规模的多语言模型在相同数据集上缩小了还原差距,表明缺陷出在标注模型本身,而非语料或翻译。主权赋予操作与性能信任;知识信任则需校准——而此审计方法成本低、可迁移至不同模型、语言与任务。

原文摘要 · Abstract (English)

National language models are becoming publicly funded epistemic infrastructure. Public ownership, linguistic specialization, and open weights create a presumption of trustworthiness. Such an instrument, built by and for a language community, looks like the natural choice for measuring what that community says and values. Whether such a model validly measures anything is untested at release. The evaluation of LLMs as measurement instruments is typically task-specific and stops at agreement with human coders. Agreement cannot distinguish an LLM instrument that measures a construct from one that reaches matching codes through surface correlates. We audit the presumption on a favourable case: AMALIA, Portugal's publicly funded 9B model, coding the moral foundation of authority in European Portuguese. The \textit{recovery gap} operationalizes the audit: decompose the codebook into its theory-defined clauses, recombine them through the theory's explicit rule, and measure how much of the original prompt's performance the stated theory reproduces. In a pre-registered, out-of-sample study on a transcreated (English to European Portuguese) corpus, AMALIA agrees with trained coders within six points of open models eight to thirteen times its size. Yet, the recovery gap shows that only about half of coding performance on authority can be attributed to the theory. A larger multilingual LLM closes the recovery gap on the same corpus, suggesting the shortfall lies in the annotator model, not the corpus or its translation. Sovereignty earns operational and performance trust; epistemic trust requires calibration -- and the audit method is inexpensive, and portable across models, languages and tasks.

大模型评估认知校准语言主权

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。