用自然语言描述人的价值观,精准预测评分差异。
Value Profiles for Encoding Human Variation
- 用上下文示范压缩成自然语言价值描述,表征个体。
- 价值描述保留超70%示范信息,优于人口统计特征。
- 可解释性强,适合个性化与标注者群体模拟。
在评分任务中建模人类差异对个性化、多元模型对齐和计算社会科学至关重要。本文提出用自然语言价值档案——从上下文示范中压缩的内在价值描述——结合可调控解码器模型,从个体表征中估计评分。我们引入信息论方法衡量表征预测能力,发现示范包含最多信息,其次为价值档案,最后是人口统计。但价值档案能有效压缩示范中的有用信息(保留>70%),并具有可解释性、可读性和可调控性优势。聚类价值档案识别相似行为个体,比最有效的统计分组更能解释评分差异。超越测试集表现,解码器预测随语义差异变化,校准良好,并可通过模拟标注者群体解释实例级分歧。结果表明,价值档案提供了超越人口统计或群体信息的新颖预测方式。
原文摘要 · Abstract (English)
Modelling human variation in rating tasks is crucial for personalization, pluralistic model alignment, and computational social science. We propose representing individuals using natural language value profiles -- descriptions of underlying values compressed from in-context demonstrations -- along with a steerable decoder model that estimates individual ratings from a rater representation. To measure the predictive information in a rater representation, we introduce an information-theoretic methodology and find that demonstrations contain the most information, followed by value profiles, then demographics. However, value profiles effectively compress the useful information from demonstrations (>70% information preservation) and offer advantages in terms of scrutability, interpretability, and steerability. Furthermore, clustering value profiles to identify similarly behaving individuals better explains rater variation than the most predictive demographic groupings. Going beyond test set performance, we show that the decoder predictions change in line with semantic profile differences, are well-calibrated, and can help explain instance-level disagreement by simulating an annotator population. These results demonstrate that value profiles offer novel, predictive ways to describe individual variation beyond demographics or group information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。