针对多语言文本嵌入,按任务类型选择性优化,提升跨任务性能。
Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation
- 按任务特性分别采用流匹配或更适配的目标函数
- 在印地语大规模文本嵌入基准上达到新最优表现
- 适合需要稳定多任务适应的多语言模型研究者
多语言文本嵌入模型通常使用单一训练目标适应多种任务,但不同任务本质需要不同的优化策略。本文提出任务条件流匹配(TCFM)框架,对翻译任务选择性应用流匹配,而对检索、分类和成对分类任务则采用更契合其学习动态的目标函数。TCFM结合教师指导的表征保留与三阶段课程学习,实现稳定适应。在Indic Massive Text Embedding Benchmark上评估,TCFM在多样化多语言任务中持续提升嵌入质量,并在不同嵌入模型族间实现良好泛化,建立新基准。论文接受后将公开代码与数据集。
原文摘要 · Abstract (English)
Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics. TCFM further combines teacher-guided representation preservation with a three-stage curriculum to enable stable adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM establishes a new state-of-the-art, consistently improving embedding quality across a diverse set of multilingual tasks and generalizing across embedding model families. We will publicly release the codebase and datasets upon acceptance of the paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。