arXiv:2603.04897cs.CL2026-03中稿 · a poster session a…

评测大模型在捕捉专家价值判断不确定性上的表现,发现其有潜力但存在偏见。

Can LLMs Capture Expert Uncertainty? A Comparative Analysis of Value Alignment in Ethnographic Qualitative Research

  • 基于舒瓦茨价值观框架,对比模型与专家对访谈中核心价值的识别
  • 模型在整体价值分布上接近人类,但排名准确率低,不确定模式不一致
  • 集成方法提升性能,但部分价值如'安全'被系统性高估

开放式访谈的定性分析在民族志和经济研究中至关重要,可揭示个体的价值观、动机及文化嵌入的金融行为。尽管大语言模型(LLMs)为自动化与深化此类解释工作带来希望,但其在任务固有模糊性下的表现仍不明确。本文评估了LLMs在识别长篇访谈中前三个舒瓦茨基本价值观(Schwartz Theory of Basic Values)方面的表现,将其输出与专家标注进行对比,分析性能与不确定性模式。结果显示,模型在集合型指标(F1、Jaccard)上接近人类上限,但在值排名准确性上表现较差(RBO得分较低)。多数模型的平均舒瓦茨价值观分布与人类分析师高度一致,但其在各价值观上的不确定性结构与专家模式存在显著差异。其中,Qwen在整体一致性上最接近专家,并最契合专家的价值分布。集成方法(如多数投票与博尔达计数)在各项指标上均有稳定提升。值得注意的是,某些价值观(如安全)被系统性高估,既显示了模型提供互补视角的潜力,也凸显需进一步探究模型引发的价值偏见。总体而言,本研究揭示了大模型在模糊性定性价值分析中的潜力与局限。

原文摘要 · Abstract (English)

Qualitative analysis of open-ended interviews plays a central role in ethnographic and economic research by uncovering individuals' values, motivations, and culturally embedded financial behaviors. While large language models (LLMs) offer promising support for automating and enriching such interpretive work, their ability to produce nuanced, reliable interpretations under inherent task ambiguity remains unclear. In our work we evaluate LLMs on the task of identifying the top three human values expressed in long-form interviews based on the Schwartz Theory of Basic Values framework. We compare their outputs to expert annotations, analyzing both performance and uncertainty patterns relative to the experts. Results show that LLMs approach the human ceiling on set-based metrics (F1, Jaccard) but struggle to recover exact value rankings, as reflected in lower RBO scores. While the average Schwartz value distributions of most models closely match those of human analysts, their uncertainty structures across the Schwartz values diverge from expert uncertainty patterns. Among the evaluated models, Qwen performs closest to expert-level agreement and exhibits the strongest alignment with expert Schwartz value distributions. LLM ensemble methods yield consistent gains across metrics, with Majority Vote and Borda Count performing best. Notably, systematic overemphasis on certain Schwartz values, like Security, suggests both the potential of LLMs to provide complementary perspectives and the need to further investigate model-induced value biases. Overall, our findings highlight both the promise and the limitations of LLMs as collaborators in inherently ambiguous qualitative value analysis.

大模型价值观分析定性研究不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。