arXiv:2508.14307cs.CL2025-08被引 2

统一标注下联合解析形态与句法,九语种表现最优。

A Joint Multitask Model for Morpho-Syntactic Parsing

  • 共享XLM-RoBERTa编码器+三任务专用解码器联合建模。
  • 平均MSLAS 78.7%,LAS 80.1%,Feats F1达90.3%。
  • 对词切分和实词识别敏感,适合多语言语法分析研究者。

我们提出一种联合多任务模型,用于UniDive 2025形态句法解析共享任务,该任务要求在新型统一树库(UD)标注体系下同时预测形态与句法分析。系统采用共享的XLM-RoBERTa编码器,并配备三个专用解码器:内容词识别、依存句法解析和形态句法特征预测。模型在涵盖九种类型差异显著语言的共享任务排行榜上取得最佳综合表现,平均MSLAS为78.7%,LAS为80.1%,特征F1达90.3%。消融实验表明,匹配任务的黄金词切分和内容词识别对性能至关重要。错误分析显示,模型在核心格范畴(尤其是主-宾格)及跨语言名词特征预测上仍存在困难。

原文摘要 · Abstract (English)

We present a joint multitask model for the UniDive 2025 Morpho-Syntactic Parsing shared task, where systems predict both morphological and syntactic analyses following novel UD annotation scheme. Our system uses a shared XLM-RoBERTa encoder with three specialized decoders for content word identification, dependency parsing, and morphosyntactic feature prediction. Our model achieves the best overall performance on the shared task's leaderboard covering nine typologically diverse languages, with an average MSLAS score of 78.7 percent, LAS of 80.1 percent, and Feats F1 of 90.3 percent. Our ablation studies show that matching the task's gold tokenization and content word identification are crucial to model performance. Error analysis reveals that our model struggles with core grammatical cases (particularly Nom-Acc) and nominal features across languages.

形态句法多任务学习多语言依存解析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。