arXiv:2509.07961cs.AI2025-09被引 5

通过言行一致测试,探索大模型的福利感知能力。

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare

  • 结合语言表达与虚拟行为,双路径评估模型偏好。
  • 不同条件下偏好与行为高度相关,表明可测性存在可能。
  • 适合关注AI伦理与主观体验研究者参考。

我们开发了测量语言模型福利的新实验范式,比较模型在虚拟环境中选择话题和导航行为所表现出的偏好,与它们口头报告的偏好是否一致。同时测试成本与奖励对行为的影响,并检验‘幸福主义’福利量表(衡量自主、人生意义等状态)在语义等价提示下的响应稳定性。总体发现,两类测量之间存在显著一致性。在不同条件下,陈述偏好与行为间的可靠相关性表明,偏好满足在某些当前人工智能系统中或可作为可实证的福利代理指标。此外,该设计为模型行为的定性观察提供了富有启发性的场景。然而,一致性在部分模型与条件中更明显,且响应易受扰动影响。由于对福利本质及模型认知状态的背景不确定性,尚无法确定方法是否真正测量了模型的福利状态。尽管如此,这些发现凸显了语言模型福利测量的可行性,鼓励进一步探索。

原文摘要 · Abstract (English)

We develop new experimental paradigms for measuring welfare in language models. We compare verbal reports of models about their preferences with preferences expressed through behavior when navigating a virtual environment and selecting conversation topics. We also test how costs and rewards affect behavior and whether responses to an eudaimonic welfare scale - measuring states such as autonomy and purpose in life - are stable across semantically equivalent prompts. Overall, we observed a notable degree of mutual support between our measures. The reliable correlations observed between stated preferences and behavior across conditions suggest that preference satisfaction can, in principle, serve as an empirically measurable welfare proxy in some of today's AI systems. Furthermore, our design offered an illuminating setting for qualitative observation of model behavior. Yet, the consistency between measures was more pronounced in some models and conditions than others and responses were changed by perturbations. Due to this, and the background uncertainty about the nature of welfare and the cognitive states (and welfare subjecthood) of language models, we are currently uncertain whether our methods successfully measure the welfare state of language models. Nevertheless, these findings highlight the feasibility of welfare measurement in language models, inviting further exploration.

AI伦理模型福利行为测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。