提出多语言时间适应框架,让分类模型更好应对跨时序数据变化。
Examining and Adapting Time for Multilingual Classification via Mixture of Temporal Experts
- 构建时间专家混合模型,融合语义与分布变化学习时间趋势。
- 实验表明不同语言在不同时期性能差异显著,MoTE有效提升泛化能力。
- 适合需要长期稳定多语言分类的系统开发者参考。
时间隐含于分类过程:分类器基于现有数据训练,却需应用于未来数据,其分布(如标签与词汇)可能随时间改变。然而现有顶尖分类模型仅关注时间变化,且主要针对英文语料,对时间效应的研究仍不充分,更缺乏多语言场景下的探索。本研究将时间视为不同域(如2024年与2025年),分析时间影响,并提出领域自适应框架以实现多语言下跨时间的分类器泛化。所提框架引入时间专家混合(Mixture of Temporal Experts, MoTE),利用语义与数据分布变化共同学习并适应时间趋势。分析显示,不同语言在不同时期表现各异;实验验证,MoTE能显著增强分类器对时间数据漂移的适应能力。研究提供分析洞见,回应了多语言场景下鲁棒时间感知模型的需求。
原文摘要 · Abstract (English)
Time is implicitly embedded in classification process: classifiers are usually built on existing data while to be applied on future data whose distributions (e.g., label and token) may change. However, existing state-of-the-art classification models merely consider the temporal variations and primarily focus on English corpora, which leaves temporal studies less explored, let alone under multilingual settings. In this study, we fill the gap by treating time as domains (e.g., 2024 vs. 2025), examining temporal effects, and developing a domain adaptation framework to generalize classifiers over time on multiple languages. Our framework proposes Mixture of Temporal Experts (MoTE) to leverage both semantic and data distributional shifts to learn and adapt temporal trends into classification models. Our analysis shows classification performance varies over time across different languages, and we experimentally demonstrate that MoTE can enhance classifier generalizability over temporal data shifts. Our study provides analytic insights and addresses the need for time-aware models that perform robustly in multilingual scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。