用互信息对齐提升多任务模型偏好,效果因任务而异
An Exploration of Self-Supervised Mutual Information Alignment for Multi-Task Settings
- 通过条件互信息引导模型响应与偏好对齐
- 单次迭代下胜过DPO 57%,数学任务表现随尝试次数提升
- 适合需要多任务适配且数据有限的场景
针对语言模型在多任务中对个体属性和偏好的精准对齐需求,本文探索自监督互信息对齐方法(SAMI)在多任务设置下的表现。首先,在多任务基准(MT-Bench)上对比SAMI与直接偏好优化(DPO),使用更强模型生成训练数据以微调更弱模型,覆盖人文、理工、抽取、编码、数学、推理和角色扮演等类别。结果表明,单次SAMI迭代对DPO的胜率为57%,不同任务间性能差异显著。其次,在GSM-8K数学任务上,相较监督微调(SFT),SAMI仅带来1.1%的零样本性能提升,而SFT达3.2%。但当允许10次尝试时,SAMI提升至3.9%,而SFT为10.1%。结合两者在多尝试场景中额外提升1.3%,单次尝试则无改善。
原文摘要 · Abstract (English)
There is a growing need for pluralistic alignment methods that can steer language models towards individual attributes and preferences. One such method, Self-Supervised Alignment with Mutual Information (SAMI), uses conditional mutual information to encourage the connection between behavioral preferences and model responses. We conduct two experiments exploring SAMI in multi-task settings. First, we compare SAMI to Direct Preference Optimization (DPO) on a multi-task benchmark (MT-Bench), using a stronger model to generate training data for a weaker one across diverse categories (humanities, STEM, extraction, coding, math, reasoning, and roleplay). Our results indicate that one iteration of SAMI has a 57% win rate against DPO, with significant variation in performance between task categories. Second, we examine SAMI's impact on mathematical accuracy (GSM-8K) relative to supervised fine-tuning (SFT). While SAMI increases zero-shot performance by 1.1%, SFT is more effective with a 3.2% boost. However, SAMI shows interesting scaling trends. When given 10 attempts, SAMI improves accuracy by 3.9%, while SFT achieves a 10.1% increase. Combining SAMI with SFT yields an additional improvement of 1.3% in multi-attempt settings, though single-attempt accuracy remains unchanged.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。