arXiv:2511.08537cs.CL2025-11

将语义角色标注数据转为情感角色标注,助力低资源场景下的情感分析。

From Semantic Roles to Opinion Roles: SRL Data Extraction for Multi-Task and Transfer Learning in Low-Resource ORL

  • 基于PropBank框架构建可复现的抽取流程,对语法树指针进行清洗与对齐。
  • 生成97,169个带角色标注的实例,精准映射至持有者-表达-目标三元组结构。
  • 成果可直接用于低资源情感角色标注任务,支持多任务与迁移学习。

本报告详述了从OntoNotes 5.0语料库中的华尔街日报部分构建高质量语义角色标注(SRL)数据集的方法,并将其适配用于情感角色标注(ORL)任务。基于PropBank标注体系,我们设计了可复现的提取流程,实现谓词-论元结构与表面文本的对齐,将句法树指针转换为连贯的文本片段,并通过严格清洗确保语义一致性。最终数据集包含97,169个谓词-论元实例,明确标注了施事(ARG0)、谓词(REL)和受事(ARG1)角色,并映射到ORL的持有者(Holder)、表达(Expression)和目标(Target)框架。文中详细描述了提取算法、非连续论元处理、标注修正及数据统计分析。该工作为研究者利用SRL提升低资源情感挖掘中的ORL性能提供了可复用资源。

原文摘要 · Abstract (English)

This report presents a detailed methodology for constructing a high-quality Semantic Role Labeling (SRL) dataset from the Wall Street Journal (WSJ) portion of the OntoNotes 5.0 corpus and adapting it for Opinion Role Labeling (ORL) tasks. Leveraging the PropBank annotation framework, we implement a reproducible extraction pipeline that aligns predicate-argument structures with surface text, converts syntactic tree pointers to coherent spans, and applies rigorous cleaning to ensure semantic fidelity. The resulting dataset comprises 97,169 predicate-argument instances with clearly defined Agent (ARG0), Predicate (REL), and Patient (ARG1) roles, mapped to ORL's Holder, Expression, and Target schema. We provide a detailed account of our extraction algorithms, discontinuous argument handling, annotation corrections, and statistical analysis of the resulting dataset. This work offers a reusable resource for researchers aiming to leverage SRL for enhancing ORL, especially in low-resource opinion mining scenarios.

情感角色标注低资源学习语义角色标注数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。