用主动学习提升大模型对少数群体的文本生成包容性。
An Active Learning Framework for Inclusive Generation by Large Language Models
- 基于聚类与知识蒸馏的主动学习框架,首次实现生成任务的有效主动学习。
- 在两个新数据集上提升基线模型2%-10%性能,跨子群体表现更稳定。
- 适合需要增强生成公平性的场景,如反叙事、风格迁移等应用。
确保大语言模型生成能代表多样子群体的文本至关重要,尤其当训练数据中少数群体相关概念稀缺时。本文提出一种基于聚类的主动学习框架,并融合知识蒸馏,首次实现生成任务的有效主动学习。该框架通过转换学习器模型的中间输出,无需预先知晓数据分布且大幅减少人工干预。通过反叙事和风格迁移的案例研究验证,构建了两个新数据集并同步训练模型,在性能上相比基线提升2%-10%。结果还显示各子群体间表现更一致,词汇多样性更高,模型对数据偏斜更具鲁棒性。此外,本方法获取的数据可提升未参与学习循环的次级模型性能,凸显框架的实际价值。
原文摘要 · Abstract (English)
Ensuring that Large Language Models (LLMs) generate text representative of diverse sub-populations is essential, particularly when key concepts related to under-represented groups are scarce in the training data. We address this challenge with a novel clustering-based active learning framework, enhanced with knowledge distillation. The proposed framework transforms the intermediate outputs of the learner model, enabling effective active learning for generative tasks for the first time. Integration of clustering and knowledge distillation yields more representative models without prior knowledge of underlying data distribution and overbearing human efforts. We validate our approach in practice through case studies in counter-narration and style transfer. We construct two new datasets in tandem with model training, showing a performance improvement of 2%-10% over baseline models. Our results also show more consistent performance across various data subgroups and increased lexical diversity, underscoring our model's resilience to skewness in available data. Further, our results show that the data acquired via our approach improves the performance of secondary models not involved in the learning loop, showcasing practical utility of the framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。