arXiv:2509.24216cs.CLcs.CY2025-09EMNLP被引 9

构建通用道德价值观分类工具,提升语言分析的可解释性与一致性

MoVa: Towards Generalizable Classification of Human Morals and Values

  • 基于四个理论框架整合16个标注数据集,支持跨领域分类
  • 轻量级提示策略超越微调模型,在多框架下表现更优
  • 适用于心理调查评估与机器行为对齐研究

识别语言中嵌入的人类道德与价值观对沟通的实证研究至关重要。然而,研究者常面临理论框架和数据资源多样化的挑战。本文提出MoVa,一个面向通用分类的人类道德与价值观资源套件,包含:(1) 16个标注数据集及来自四个理论基础框架的基准结果;(2) 一种轻量级大模型提示策略,在多个领域和框架中表现优于微调模型;(3) 一项新应用,用于评估心理问卷。实践中推荐使用all@once策略,一次性评分所有相关概念,类似多标签分类链。MoVa的数据与方法可促进人类与机器沟通的细粒度解读,对机器行为对齐具有潜在意义。

原文摘要 · Abstract (English)

Identifying human morals and values embedded in language is essential to empirical studies of communication. However, researchers often face substantial difficulty navigating the diversity of theoretical frameworks and data available for their analysis. Here, we contribute MoVa, a well-documented suite of resources for generalizable classification of human morals and values, consisting of (1) 16 labeled datasets and benchmarking results from four theoretically-grounded frameworks; (2) a lightweight LLM prompting strategy that outperforms fine-tuned models across multiple domains and frameworks; and (3) a new application that helps evaluate psychological surveys. In practice, we specifically recommend a classification strategy, all@once, that scores all related concepts simultaneously, resembling the well-known multi-label classifier chain. The data and methods in MoVa can facilitate many fine-grained interpretations of human and machine communication, with potential implications for the alignment of machine behavior.

道德识别大模型提示价值观分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。