arXiv:2606.04189cs.CL2026-06中稿 · The 28th Internati…

ACAT让多人协作标注情感分析数据更高效,自动整合结果并计算一致性。

ACAT: A Collaborative Platform for Efficient Aspect-Based Sentiment Dataset Annotation

  • 基于网页的协同标注平台,原生支持四种情感分析任务。
  • 自动处理多标注数据,导出即得可用训练集,平均单条标注31.58秒。
  • 适合需要高质量标注数据的研究者,尤其适合团队协作场景。

基于方面的情感分析(ABSA)需要高质量数据集以训练可靠模型。然而现有标注工具将输出视为扁平文件,研究人员需手动合并多标注数据、重构关系结构,并通过自定义脚本计算可靠性指标。本文提出ACAT(面向方面情感分析的协同标注工具),一个原生支持四种ABSA工作流的网络平台:(1) 方面类别情感分析,(2) 句子级分割,(3) 带字符级位置追踪的方面项情感分析,(4) 保留双重跨度偏移的方面情感三元组抽取。其核心贡献是自动化提取、转换、加载(ETL)流程,可对协同标注进行对齐,并在导出时直接计算标注者间一致性(IAA)度量,生成可直接用于训练的数据集。在包含1,002条餐厅评论的初步验证中,两名不同专业背景的标注员完成任务,平均标注时间31.58秒,各任务原始IAA值介于0.78至0.86之间。

原文摘要 · Abstract (English)

Aspect-Based Sentiment Analysis (ABSA) requires high-quality datasets to train reliable models. However, existing annotation tools treat output as flat files, leaving researchers to manually consolidate multi-annotator data, reconstruct relational structures, and compute reliability metrics through custom scripts. This paper introduces ACAT (Aspect-based sentiment analysis Collaborative Annotation Tool), a web-based platform natively supporting four ABSA workflows: (1) Aspect-Category Sentiment Analysis, (2) Clause-Level Segmentation, (3) Aspect-Term Sentiment Analysis with character-level position tracking, and (4) Aspect Sentiment Triplet Extraction with dual span offset preservation. Its core contribution is an automated Extract, Transform, Load (ETL) pipeline that aligns collaborative annotations and computes Inter-Annotator Agreement (IAA) metrics directly at export, yielding training-ready datasets. In a preliminary validation on 1,002 restaurant reviews with two annotators of differing expertise, ACAT achieves a median annotation time of 31.58 seconds and a raw IAA ranging from 0.78 to 0.86 across all tasks.

情感分析数据标注协同工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。