多语言训练让写作风格模型跨语言更准,提升泛化能力。
Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domains
- 用概率掩码让模型关注风格词而非内容词
- 跨语言批量处理减少干扰,提升对比学习效果
- 在21种非英语语言上平均召回率提升4.85%
作者风格表示(AR)学习通过建模作者独特写作风格,在作者归属任务中表现优异。但以往研究主要集中于单语场景(尤其是英语),多语言AR模型的潜力未被充分探索。本文提出一种新型多语言AR学习方法,包含两项关键创新:概率内容掩码,促使模型聚焦风格指示性词汇而非内容特异性词汇;语言感知批处理,通过减少跨语言干扰优化对比学习。模型在超过450万作者、36种语言和13个领域上进行训练。在22种非英语语言中的21种上,性能持续优于单语基线,平均召回率@8提升4.85%,单语言最高提升达15.91%。此外,相比仅用英语训练的单语模型,该模型展现出更强的跨语言与跨领域泛化能力。分析证实两项技术均有效,对性能提升至关重要。
原文摘要 · Abstract (English)
Authorship representation (AR) learning, which models an author's unique writing style, has demonstrated strong performance in authorship attribution tasks. However, prior research has primarily focused on monolingual settings-mostly in English-leaving the potential benefits of multilingual AR models underexplored. We introduce a novel method for multilingual AR learning that incorporates two key innovations: probabilistic content masking, which encourages the model to focus on stylistically indicative words rather than content-specific words, and language-aware batching, which improves contrastive learning by reducing cross-lingual interference. Our model is trained on over 4.5 million authors across 36 languages and 13 domains. It consistently outperforms monolingual baselines in 21 out of 22 non-English languages, achieving an average Recall@8 improvement of 4.85%, with a maximum gain of 15.91% in a single language. Furthermore, it exhibits stronger cross-lingual and cross-domain generalization compared to a monolingual model trained solely on English. Our analysis confirms the effectiveness of both proposed techniques, highlighting their critical roles in the model's improved performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。