arXiv:2509.11498cs.CL2025-09被引 5

基于MT5和Qwen的多模型系统,提升低资源语言论述关系分类性能。

DeDisCo at the DISRPT 2025 Shared Task: A System for Discourse Relation Classification

  • 融合MT5编码器与Qwen解码器,结合自动翻译增强数据。
  • 在跨语言任务中达到71.28%的宏平均准确率。
  • 适用于低资源语言论述分析,适合自然语言处理研究者。

本文介绍德吉大学在DISRPT 2025共享任务中的论述关系分类系统DeDisCo。我们测试了两种方法:基于mt5的编码器与基于开源Qwen模型的解码器。同时,针对低资源语言,通过自动将英文匹配数据翻译生成增广数据集,并引入前几届共享任务中表现优异的语法特征。该系统在评测中取得71.28%的宏平均准确率,并对结果进行了可解释性分析与错误归因。

原文摘要 · Abstract (English)

This paper presents DeDisCo, Georgetown University's entry in the DISRPT 2025 shared task on discourse relation classification. We test two approaches, using an mt5-based encoder and a decoder based approach using the openly available Qwen model. We also experiment on training with augmented dataset for low-resource languages using matched data translated automatically from English, as well as using some additional linguistic features inspired by entries in previous editions of the Shared Task. Our system achieves a macro-accuracy score of 71.28, and we provide some interpretation and error analysis for our results.

论述关系低资源多语言MT5

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。