用主成分分析生成初始偏好标签,解决冷启动难题。
Cold-Start Active Preference Learning in Socio-Economic Domains
- 用PCA从数据结构中自监督生成初始伪标签。
- 在金融信用、职业成功率等数据集上显著优于无先验方法。
- 适合缺乏初始标注的社科领域偏好建模场景。
主动偏好学习虽能高效建模偏好,但冷启动问题导致无初始标注时性能骤降。现有解决方案多集中于视觉与文本领域,而社会经济领域的冷启动问题尚未深入探索。本文受社会经济研究启发,提出先通过主成分分析(PCA)生成初始伪标签,实现无需专家输入的自监督预热,构建初步模型;随后进入主动学习循环,向模拟噪声代理查询标签以优化模型。在金融信用、职业成功率及社会经济地位等多个真实数据集上的实验表明,该方法显著优于传统无先验主动学习策略,提供了一种计算高效且易实现的冷启动解决方案。
原文摘要 · Abstract (English)
Active preference learning offers an efficient approach to modeling preferences, but it is hindered by the cold-start problem, which leads to a marked decline in performance when no initial labeled data are available. While cold-start solutions have been proposed for domains such as vision and text, the cold-start problem in active preference learning remains largely unexplored, underscoring the need for practical, effective methods. Drawing inspiration from established practices in social and economic research, the proposed method initiates learning with a self-supervised phase that employs Principal Component Analysis (PCA) to generate initial pseudo-labels. This process produces a \say{warmed-up} model based solely on the data's intrinsic structure, without requiring expert input. The model is then refined through an active learning loop that strategically queries a simulated noisy oracle for labels. Experiments conducted on various socio-economic datasets, including those related to financial credibility, career success rate, and socio-economic status, consistently show that the PCA-driven approach outperforms standard active learning strategies that start without prior information. This work thus provides a computationally efficient and straightforward solution that effectively addresses the cold-start problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。