arXiv:2502.20552cs.CL2025-02

首个匈牙利语抽象意义表示数据集及解析器,助力非英语语义分析研究。

HuAMR: A Hungarian AMR Parser and Dataset

  • 用大模型自动生成银标准标注,人工精修构建高质量匈牙利语数据集。
  • 在新闻文本上,小模型结合银标准数据后解析准确率显著提升。
  • 适合从事低资源语言语义解析、多语言NLP的研究者参考。

我们提出HuAMR,首个面向匈牙利语的抽象意义表示(AMR)数据集及基于大语言模型的AMR解析器,旨在缓解非英语语言语义资源稀缺的问题。为构建HuAMR,我们利用Llama-3.1-70B模型自动生成银标准AMR标注,并通过人工校对确保质量。基于该数据集,我们评估了mT5 Large与Llama-3.2-1B两种模型架构及微调策略对解析性能的影响。尽管将Llama-3.1-70B生成的银标准数据纳入小模型训练并未一致提升整体得分,但结果显示其显著提升了在匈牙利语新闻数据上的解析准确率(目标领域)。我们采用Smatch评分进行评估,验证了HuAMR及所提解析器在推动语义解析研究方面的潜力。

原文摘要 · Abstract (English)

We present HuAMR, the first Abstract Meaning Representation (AMR) dataset and a suite of large language model-based AMR parsers for Hungarian, targeting the scarcity of semantic resources for non-English languages. To create HuAMR, we employed Llama-3.1-70B to automatically generate silver-standard AMR annotations, which we then refined manually to ensure quality. Building on this dataset, we investigate how different model architectures - mT5 Large and Llama-3.2-1B - and fine-tuning strategies affect AMR parsing performance. While incorporating silver-standard AMRs from Llama-3.1-70B into the training data of smaller models does not consistently boost overall scores, our results show that these techniques effectively enhance parsing accuracy on Hungarian news data (the domain of HuAMR). We evaluate our parsers using Smatch scores and confirm the potential of HuAMR and our parsers for advancing semantic parsing research.

语义解析匈牙利语大模型AMR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。