提出衡量大模型语言人性化程度的方法,发现用户更偏好非人性化输出。
HumT DumT: Measuring and controlling human-like language in LLMs
- 基于大模型自身概率计算人性化度量(HumT)
- 用户在多数场景更偏好非人性化语言输出
- 可控制降低语言人性化程度,适合关注伦理风险的开发者
大模型生成类人语言可能提升用户体验,但也可能引发欺骗、过度依赖和刻板印象。本文提出HumT与SocioT两个度量指标,基于大模型相对概率评估文本的人性化程度及社会感知维度。在偏好和使用数据集上测量显示,用户在多数情境中更偏好非人性化输出。HumT揭示类人语言与温暖感、社交亲近、女性化、低地位高度相关,这些特征与前述风险紧密关联。为此,本文提出DumT方法,利用HumT系统性控制并降低语言人性化程度,同时保持模型性能。该方法为缓解类人语言生成带来的风险提供了实用路径。
原文摘要 · Abstract (English)
Should LLMs generate language that makes them seem human? Human-like language might improve user experience, but might also lead to deception, overreliance, and stereotyping. Assessing these potential impacts requires a systematic way to measure human-like tone in LLM outputs. We introduce HumT and SocioT, metrics for human-like tone and other dimensions of social perceptions in text data based on relative probabilities from an LLM. By measuring HumT across preference and usage datasets, we find that users prefer less human-like outputs from LLMs in many contexts. HumT also offers insights into the perceptions and impacts of anthropomorphism: human-like LLM outputs are highly correlated with warmth, social closeness, femininity, and low status, which are closely linked to the aforementioned harms. We introduce DumT, a method using HumT to systematically control and reduce the degree of human-like tone while preserving model performance. DumT offers a practical approach for mitigating risks associated with anthropomorphic language generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。