arXiv:2602.22865cs.CL2026-02Conference of the …

用问答式框架跨语言自动标注谓词-论元关系,提升多语言语义分析效率。

Effective QA-driven Annotation of Predicate-Argument Relations Across Languages

  • 以问答形式重构谓词-论元关系,通过翻译与对齐实现跨语言迁移。
  • 在希伯来语、俄语、法语上生成高质量训练数据,模型性能超越主流大模型。
  • 无需从零标注,适合需要快速构建多语言语义系统的研究者。

显式的谓词-论元关系表征是可解释语义分析的基础,支持推理、生成与评估。然而,获取此类语义结构需大量人工标注,且长期局限于英语。本文利用问答驱动的语义角色标注(QA-SRL)框架——一种自然语言表述的谓词-论元关系形式——作为跨语言扩展语义标注的基础。为此,提出一种跨语言投影方法,通过受控翻译与词对齐流程,复用英语QA-SRL解析器,自动生成与目标语言谓词对齐的问答标注。该方法应用于希伯来语、俄语和法语(涵盖不同语系),生成高质量训练数据,并微调出性能优于强大多语言大模型(GPT-4o、LLaMA-Maverick)的语言特异性解析器。通过将QA-SRL作为可迁移的自然语言语义接口,本方法实现了高效、广泛可用的跨语言谓词-论元解析。

原文摘要 · Abstract (English)

Explicit representations of predicate-argument relations form the basis of interpretable semantic analysis, supporting reasoning, generation, and evaluation. However, attaining such semantic structures requires costly annotation efforts and has remained largely confined to English. We leverage the Question-Answer driven Semantic Role Labeling (QA-SRL) framework -- a natural-language formulation of predicate-argument relations -- as the foundation for extending semantic annotation to new languages. To this end, we introduce a cross-linguistic projection approach that reuses an English QA-SRL parser within a constrained translation and word-alignment pipeline to automatically generate question-answer annotations aligned with target-language predicates. Applied to Hebrew, Russian, and French -- spanning diverse language families -- the method yields high-quality training data and fine-tuned, language-specific parsers that outperform strong multilingual LLM baselines (GPT-4o, LLaMA-Maverick). By leveraging QA-SRL as a transferable natural-language interface for semantics, our approach enables efficient and broadly accessible predicate-argument parsing across languages.

语义角色标注跨语言问答系统自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。