将偏好模型的决策原因用自然语言解释并可编辑,让机器判断更透明可信。
From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language

- 自动提取偏好维度并用自然语言描述,关联模型内部表示。
- 在450人道德困境和449人电影选择实验中提升预测准确率。
- 支持用户实时检查与修改模型推理,适合需可解释性的应用。
统计学习算法从高维选择数据中推断人类偏好时面临核心挑战:选择项通常同时存在多个差异,难以确定哪些因素真正驱动了决策。此外,这些方法的黑箱特性使操作者无法检查、质疑或纠正错误。本文提出“权重到文字”方法,输入选择数据集后,自动发现一组领域相关的偏好维度,每个维度以自然语言描述,并对应模型表示空间中的向量。该方法解决判定模糊与不透明问题:可聚焦于少数有意义因素进行归因,且将模型推断外化为自然语言,供用户实时审查与修改。我们先在道德困境、电影、葡萄酒及自由形式大模型输出四个领域定性展示其通用性;再通过两项预注册的人类实验(道德困境,N=450;电影选择,N=449)验证其优势:(1) 将偏好模型正则化至所学基底能提高对保留选择的预测准确率;(2) 引入参与者结构化修改后准确率进一步提升。一对一比较中,参与者更倾向该方法推导出的偏好画像,并认为其预测更准确。
原文摘要 · Abstract (English)
The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is generally unclear which factors actually drove an observed decision and should be credited as preferences. Compounding this problem, the opacity of these methods leaves human operators unable to inspect, contest, or correct models when they err. We introduce \emph{weights to words}, a method that takes a dataset of choice problems as input and automatically discovers a collection of domain-relevant preference dimensions, each described in natural language and paired with a vector in the model's representational space. These dimensions address both under-determination and opacity: they can be applied to concentrate attribution on a small set of meaningful factors, and they can externalize the model's inferences in natural language so that users can inspect and edit them in real time. We first qualitatively illustrate the method's versatility on four diverse domains: moral dilemmas, movies, wines, and free-form LLM responses. We then report two pre-registered human-subjects experiments, on moral dilemmas ($N=450$) and movie selection ($N=449$), that demonstrate its benefits for learning preference models: (1) regularizing a preference model toward the learned basis increases prediction accuracy on held-out choices, and (2) incorporating participants' structured edits further improves accuracy. In head-to-head comparisons, participants prefer the method's inferred preference profiles and endorse its predictions as more accurate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。