用代码本引导对话分段,提升标注一致性。
Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation
- 基于下游任务需求设计分段边界,增强语义连贯性。
- 大模型分段更一致,但全局转折检测仍逊于传统方法。
- 适合需要精准对话结构的任务,如教育对话分析。
对话行为标注通常将交际或教学意图视为局限于单个话语或回合,导致标注者对底层行为达成一致,却在分段边界上分歧,降低可靠性。我们提出代码本注入的分段方法,使边界判断基于下游标注标准,并对比基于大模型与标准及检索增强基线的分段器。为无真实标签评估,引入跨度一致性、区分度及人机分布一致性指标。结果表明,具备对话行为意识的分段器内部一致性更高;大模型擅长生成任务一致的片段,但基于连贯性的基线在捕捉对话流全局变化方面仍更优。在两个数据集上,无单一分段器全面领先。段内连贯性提升常伴随边界区分度与人机分布一致性下降。研究强调分段是关键设计选择,应针对下游目标优化,而非单一性能指标。
原文摘要 · Abstract (English)
Dialogue Act (DA) annotation typically treats communicative or pedagogical intent as localized to individual utterances or turns. This leads annotators to agree on the underlying action while disagreeing on segment boundaries, reducing apparent reliability. We propose codebook-injected segmentation, which conditions boundary decisions on downstream annotation criteria, and evaluate LLM-based segmenters against standard and retrieval-augmented baselines. To assess these without gold labels, we introduce evaluation metrics for span consistency, distinctiveness, and human-AI distributional agreement. We found DA-awareness produces segments that are internally more consistent than text-only baselines. While LLMs excel at creating construct-consistent spans, coherence-based baselines remain superior at detecting global shifts in dialogue flow. Across two datasets, no single segmenter dominates. Improvements in within-segment coherence frequently trade off against boundary distinctiveness and human-AI distributional agreement. These results highlight segmentation as a consequential design choice that should be optimized for downstream objectives rather than a single performance score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。