arXiv:2504.08961cs.CL2025-04ACL被引 2

用大模型自动构建对话篇章标注体系并标注,效率远超人工。

A Fully Automated Pipeline for Conversational Discourse Annotation: Tree Scheme Generation and Labeling with Large Language Models

  • 用大模型自动生成树状标注框架,无需人工设计。
  • 频率引导的决策树搭配先进大模型,效果超越人工标注。
  • 适合想快速开展对话分析的研究者使用。

大语言模型在自动化对话篇章标注方面展现出巨大潜力。尽管手动设计树状标注体系能显著提升标注质量,但其构建过程耗时且需专业知识。本文提出一个完全自动化的流水线,利用大模型生成标注体系并完成标注。我们在语用功能(SFs)和Switchboard-DAMSL(SWBD-DAMSL)分类体系上评估该方法,对比了多种设计选择。实验表明,基于频率引导的决策树配合先进大模型进行标注,性能可超越此前人工设计的体系,甚至达到或超过人类标注者水平,同时大幅缩短标注时间。我们已公开全部代码、生成的标注体系及标注结果,以促进后续对话篇章标注研究。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) have shown promise in automating discourse annotation for conversations. While manually designing tree annotation schemes significantly improves annotation quality for humans and models, their creation remains time-consuming and requires expert knowledge. We propose a fully automated pipeline that uses LLMs to construct such schemes and perform annotation. We evaluate our approach on speech functions (SFs) and the Switchboard-DAMSL (SWBD-DAMSL) taxonomies. Our experiments compare various design choices, and we show that frequency-guided decision trees, paired with an advanced LLM for annotation, can outperform previously manually designed trees and even match or surpass human annotators while significantly reducing the time required for annotation. We release all code and resultant schemes and annotations to facilitate future research on discourse annotation.

对话标注大模型应用自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。