arXiv:2507.12041cs.LG2025-07

更细粒度的反馈能显著提升小样本预测效果

Granular feedback merits sophisticated aggregation

  • 用复杂聚合方法替代简单平均,提升反馈预测精度
  • 五点反馈下,只需一半人数就能达到相同效果
  • 适合小样本场景下的人类反馈分析任务

人类反馈在训练AI模型、推荐系统和民意测量中日益重要,细粒度反馈因信息量更大而优于二元反馈。尽管大规模群体反馈可准确估计总体分布,但成本限制常需使用小样本。现有常用方法是正则化平均:基于经验分布并朝先验正则化。本文发现,随着反馈粒度增加,通过更复杂的个体反馈组合方式,可显著超越正则化平均的预测性能。实证分析基于社会态度调查问题验证了这一规律:二元反馈下,提升方法复杂性几乎不减少所需人数;而五点反馈时,复杂方法仅需约一半人数即可达到正则化平均的性能水平。

原文摘要 · Abstract (English)

Human feedback is increasingly used across diverse applications like training AI models, developing recommender systems, and measuring public opinion -- with granular feedback often being preferred over binary feedback for its greater informativeness. While it is easy to accurately estimate a population's distribution of feedback given feedback from a large number of individuals, cost constraints typically necessitate using smaller groups. A simple method to approximate the population distribution is regularized averaging: compute the empirical distribution and regularize it toward a prior. Can we do better? As we will discuss, the answer to this question depends on feedback granularity. Suppose one wants to predict a population's distribution of feedback using feedback from a limited number of individuals. We show that, as feedback granularity increases, one can substantially improve upon predictions of regularized averaging by combining individuals' feedback in ways more sophisticated than regularized averaging. Our empirical analysis using questions on social attitudes confirms this pattern. In particular, with binary feedback, sophistication barely reduces the number of individuals required to attain a fixed level of performance. By contrast, with five-point feedback, sophisticated methods match the performance of regularized averaging with about half as many individuals.

人类反馈小样本学习反馈聚合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。