基于MT5和Qwen的多模型系统,提升低资源语言论述关系分类性能。
DeDisCo at the DISRPT 2025 Shared Task: A System for Discourse Relation Classification
- 融合MT5编码器与Qwen解码器,结合自动翻译增强数据。
- 在跨语言任务中达到71.28%的宏平均准确率。
- 适用于低资源语言论述分析,适合自然语言处理研究者。
本文介绍德吉大学在DISRPT 2025共享任务中的论述关系分类系统DeDisCo。我们测试了两种方法:基于mt5的编码器与基于开源Qwen模型的解码器。同时,针对低资源语言,通过自动将英文匹配数据翻译生成增广数据集,并引入前几届共享任务中表现优异的语法特征。该系统在评测中取得71.28%的宏平均准确率,并对结果进行了可解释性分析与错误归因。
原文摘要 · Abstract (English)
This paper presents DeDisCo, Georgetown University's entry in the DISRPT 2025 shared task on discourse relation classification. We test two approaches, using an mt5-based encoder and a decoder based approach using the openly available Qwen model. We also experiment on training with augmented dataset for low-resource languages using matched data translated automatically from English, as well as using some additional linguistic features inspired by entries in previous editions of the Shared Task. Our system achieves a macro-accuracy score of 71.28, and we provide some interpretation and error analysis for our results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。