arXiv:2503.16480cs.HCcs.AI2025-03被引 2

研究用户如何偏好有理有据的对话,揭示AI对人类互动风格的模仿机制。

Human Preferences for Constructive Interactions in Language Model Alignment

  • 基于74国超7500条对话数据,分析人类对语言模型回应的偏好
  • 用户更倾向有逻辑、有深度的回应,而非个人故事化表达
  • 用户可主动引导对话基调,模型会镜像其语言特征(如毒性)

随着大型语言模型(LLMs)进入主流,使其促进建设性对话而非加剧社会分裂至关重要。我们利用涵盖74个国家、超过7500次对话的个性化多元文化对话语料库,考察了与建设性互动相关的语言特征在用于训练AI的人类偏好数据中的体现。研究发现,用户始终更偏好逻辑清晰、观点细腻的回应,而排斥高度个人叙事化的表达。然而,认为AI应反映自身价值观的用户,对推理的要求较低,反而更看重好奇心。令人鼓舞的是,用户能主动设定对话的建设性程度,因为语言模型会镜像用户提问中的语言特征,包括毒性水平。

原文摘要 · Abstract (English)

As large language models (LLMs) enter the mainstream, aligning them to foster constructive dialogue rather than exacerbate societal divisions is critical. Using an individualized and multicultural alignment dataset of over 7,500 conversations of individuals from 74 countries engaging with 21 LLMs, we examined how linguistic attributes linked to constructive interactions are reflected in human preference data used for training AI. We found that users consistently preferred well-reasoned and nuanced responses while rejecting those high in personal storytelling. However, users who believed that AI should reflect their values tended to place less preference on reasoning in LLM responses and more on curiosity. Encouragingly, we observed that users could set the tone for how constructive their conversation would be, as LLMs mirrored linguistic attributes, including toxicity, in user queries.

语言模型对齐用户偏好建设性对话多文化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。