通过元学习强化评分者示例,提升模型对主观任务的适应能力
Opt-ICL at LeWiDi-2025: Maximizing In-Context Signal from Rater Examples via Meta-Learning
- 采用两阶段元学习:先在多数据集上后训练,再针对特定数据分布进行上下文元学习
- 在LeWiDi竞赛中两项任务均夺冠,关键组件实验表明评分者示例不可或缺
- 适合处理标注分歧大的主观类NLP任务,如评价、情感分析等场景
许多自然语言处理任务涉及主观性、模糊性或标注者间的合理分歧。本文提出一种建模人类差异的系统,利用大语言模型的上下文学习能力,并通过两阶段元学习训练:1)在多个需上下文学习的数据集上进行后训练;2)通过上下文元学习使模型适配目标数据分布。该系统在“学习分歧”(LeWiDi)竞赛中两项任务均获冠军。我们还进行了消融实验,发现:包含评分者示例在上下文中对性能至关重要,针对大尺寸数据集进行特定微调有帮助,在一个竞赛数据集上使用其他上下文数据集后训练也有效,且模型规模越大性能越优。
原文摘要 · Abstract (English)
Many natural language processing (NLP) tasks involve subjectivity, ambiguity, or legitimate disagreement between annotators. In this paper, we outline our system for modeling human variation. Our system leverages language models' (LLMs) in-context learning abilities, along with a two-step meta-learning training procedure for 1) post-training on many datasets requiring in-context learning and 2) specializing the model via in-context meta-learning to the particular data distribution of interest. We also evaluate the performance of our system submission to the Learning With Disagreements (LeWiDi) competition, where it was the overall winner on both tasks. Additionally, we perform an ablation study to measure the importance of each system component. We find that including rater examples in-context is crucial for our system's performance, dataset-specific fine-tuning is helpful on the larger datasets, post-training on other in-context datasets is helpful on one of the competition datasets, and that performance improves with model scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。