用专家关联显式校准文本嵌入,让模型理解更贴近人类意图。
Grounding Text Embeddings in Stakeholder Associations

- 通过专家关联实验显性化领域知识,建立人机语义对齐标准。
- 模型与专家在丹麦政策分析中存在19-26个百分点的语义距离差距。
- 该方法可复用于不同语言与领域,适合需可信语义分析的研究者。
文本嵌入广泛用于分析复杂文本语料,但其捕捉的语义距离是否与人类专家一致尚不明确。确保嵌入表示与人类意图对齐,是保证分析有效性的关键。本文提出「利益相关者校准实验」,旨在显式表达专家关联,并将嵌入模型结果与人类理解对齐。在丹麦政策议题的主案例研究中,发现神经文本嵌入的可靠性显著低于人类专家(19-26个百分点差距),且该偏差会传递至下游聚类性能(作业排名与聚类质量间斯皮尔曼相关系数ρ=0.9)。对美国联邦人工智能应用的次级研究在英文语境下复现了16个百分点的差距,采用数字协议与另一专家群体,证明该差距非单一工具或领域的偶然现象。该方法为评估嵌入模型是否捕捉领域专家最关心的语义差异提供了实用路径。
原文摘要 · Abstract (English)
Text embeddings are widely used to analyse large corpora of complex texts. However, it is unclear whether the embeddings capture the same semantic distances as the human experts using them. Ensuring alignment between embedding representations and human intentions is essential for valid analyses. We present the Stakeholder Grounding Exercise, a method for making expert associations explicit and grounding embedding model results in human understanding. In our primary case study on Danish policy issues, we find that neural text embeddings are substantially less reliable than human experts (19-26 pp gap), and that this misalignment propagates to downstream clustering performance (Spearman $ρ=0.9$ between exercise ranking and cluster quality). A secondary study on US Federal AI use cases replicates the gap (16pp) in English, using a digital protocol and a different community of experts -- demonstrating that the gap is not an artefact of a single instrument or domain. The Stakeholder Grounding Exercise offers a practical method for assessing whether embedding models capture the semantic distinctions that matter most to domain experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。