arXiv:2412.04936cs.CLcs.AI2024-12被引 8

对比文本、行为与脑数据的语义表征,发现行为数据能捕捉独特情感与道德维度。

Probing the contents of semantic representations from text, behavior, and brain data using the psychNorms metabase

  • 用表征相似性分析比较文本、行为和脑数据的语义表征差异。
  • 行为表征在情感、主动性和社会道德维度上包含文本未覆盖的独特信息。
  • 适用于研究人类对齐语义表征,尤其对大模型对齐评估有参考价值。

语义表征在自然语言处理、心理语言学和人工智能中至关重要。尽管传统上依赖互联网文本生成,近年来基于行为(如自由联想)和脑数据(如fMRI)的表征日益流行,有望提升对人类表征的测量与建模能力。本文首次系统评估了源自文本、行为和脑数据的语义表征之间的异同。通过表征相似性分析,我们发现行为与脑数据导出的词向量所编码的信息不同于文本衍生的版本。进一步利用我们的psychNorms metabase及一种称为表征内容分析的可解释性方法,我们发现行为表征在特定情感、主动性及社会道德维度上捕获了独特的方差。因此,行为数据成为补充文本以捕捉人类语义表征的重要来源。这些结果广泛适用于旨在学习人类对齐语义表征的研究,包括大型语言模型的评估与对齐工作。

原文摘要 · Abstract (English)

Semantic representations are integral to natural language processing, psycholinguistics, and artificial intelligence. Although often derived from internet text, recent years have seen a rise in the popularity of behavior-based (e.g., free associations) and brain-based (e.g., fMRI) representations, which promise improvements in our ability to measure and model human representations. We carry out the first systematic evaluation of the similarities and differences between semantic representations derived from text, behavior, and brain data. Using representational similarity analysis, we show that word vectors derived from behavior and brain data encode information that differs from their text-derived cousins. Furthermore, drawing on our psychNorms metabase, alongside an interpretability method that we call representational content analysis, we find that, in particular, behavior representations capture unique variance on certain affective, agentic, and socio-moral dimensions. We thus establish behavior as an important complement to text for capturing human representations and behavior. These results are broadly relevant to research aimed at learning human-aligned semantic representations, including work on evaluating and aligning large language models.

语义表征行为数据脑数据大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。