arXiv:2410.03357cs.CL2024-10EMNLP被引 5

对比元学习与联合学习在跨语言语义解析中的效果,发现0样本下元学习略优。

Should Cross-Lingual AMR Parsing go Meta? An Empirical Assessment of Meta-Learning and Joint Learning AMR Parsing

  • 采用元学习框架,模拟小样本快速适应新语言的解析任务
  • 0样本时部分语言性能略胜传统联合学习,但超1样本后优势消失
  • 首次为克罗地亚语和韩语构建公开可用的语义解析数据集

跨语言语义解析(AMR)指在仅拥有源语言训练数据的情况下,为目标语言生成语义图。由于AMR数据量小,该任务此前仅在英语、西班牙语、德语、中文和意大利语等少数语言中探索。受Langedijk等人(2022)启发,我们研究将元学习应用于跨语言语义解析。我们在k-shot(含0-shot)场景下评估模型在克罗地亚语、波斯语、韩语、中文和法语上的表现。值得注意的是,韩语和克罗地亚语测试集基于现有的《小王子》英文语料库构建,并已公开。通过与经典联合学习方法对比,实验表明:尽管元学习模型在0样本条件下对某些语言有轻微提升,但在k > 0时性能增益微弱或不显著。

原文摘要 · Abstract (English)

Cross-lingual AMR parsing is the task of predicting AMR graphs in a target language when training data is available only in a source language. Due to the small size of AMR training data and evaluation data, cross-lingual AMR parsing has only been explored in a small set of languages such as English, Spanish, German, Chinese, and Italian. Taking inspiration from Langedijk et al. (2022), who apply meta-learning to tackle cross-lingual syntactic parsing, we investigate the use of meta-learning for cross-lingual AMR parsing. We evaluate our models in $k$-shot scenarios (including 0-shot) and assess their effectiveness in Croatian, Farsi, Korean, Chinese, and French. Notably, Korean and Croatian test sets are developed as part of our work, based on the existing The Little Prince English AMR corpus, and made publicly available. We empirically study our method by comparing it to classical joint learning. Our findings suggest that while the meta-learning model performs slightly better in 0-shot evaluation for certain languages, the performance gain is minimal or absent when $k$ is higher than 0.

跨语言解析元学习语义分析小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。