首个多语言立场抽取基准,支持六种语言立场分析
Multilingual Target-Stance Extraction
- 构建跨语言通用框架,无需每种语言单独建模
- 多语言环境下F1仅12.78,目标识别成主要瓶颈
- 首次揭示目标表述方式对结果的显著影响
社交媒体为争议性议题的公众意见数据驱动分析提供了可能。目标立场抽取(TSE)旨在识别文本讨论的目标及其立场。现有研究多聚焦英语场景,而本文首次提出覆盖加泰罗尼亚语、爱沙尼亚语、法语、意大利语、中文和西班牙语的多语言TSE基准。通过统一模型架构实现跨语言泛化,无需为每种语言单独训练模型。当前方法在多语言设置下仅取得12.78的F1分数,显著低于英语单语场景,凸显目标识别的困难。同时,本文首次验证了不同目标表述形式对F1得分的敏感性。这些成果为多语言TSE提供关键资源与评估基准。
原文摘要 · Abstract (English)
Social media enables data-driven analysis of public opinion on contested issues. Target-Stance Extraction (TSE) is the task of identifying the target discussed in a document and the document's stance towards that target. Many works classify stance towards a given target in a multilingual setting, but all prior work in TSE is English-only. This work introduces the first multilingual TSE benchmark, spanning Catalan, Estonian, French, Italian, Mandarin, and Spanish corpora. It manages to extend the original TSE pipeline to a multilingual setting without requiring separate models for each language. Our model pipeline achieves a modest F1 score of 12.78, underscoring the increased difficulty of the multilingual task relative to English-only setups and highlighting target prediction as the primary bottleneck. We are also the first to demonstrate the sensitivity of TSE's F1 score to different target verbalizations. Together these serve as a much-needed baseline for resources, algorithms, and evaluation criteria in multilingual TSE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。