提出新评估框架,解决对话分段中边界精度误判问题
When F1 Fails: Granularity-Aware Evaluation for Dialogue Topic Segmentation
- 分离边界评分与选择,按粒度区间评估分段质量
- 发现不同数据集性能差异多由标注粒度不一致导致
- 适合需灵活处理对话上下文的LLM系统开发者
对话分段支持摘要生成、信息检索、记忆管理与对话连贯性。尽管研究多年,评估仍以严格边界匹配和F1指标为主。现代大语言模型对话系统依赖分段来管理超出固定上下文窗口的对话历史,而未结构化的内容积累会降低效率与连贯性。本文提出一种新评估框架,报告边界密度及段落对齐诊断(纯度与覆盖率)并引入窗口容忍F1(W-F1)。通过分离边界评分与选择,可在不同粒度范围内评估分段质量,而非单一操作点。跨数据集评估显示,性能差异常源于标注粒度不匹配,而非边界定位质量本身。在八组涵盖任务导向、开放域、会议式与合成交互的对话数据集上,评估了结构不同的分段策略。边界指标强依赖边界密度:阈值扫描带来的W-F1变化大于方法切换。这表明分段应视为粒度选择问题,而非预测唯一正确边界集。因此,建议分离边界评分与选择,以分析和调优不同标注粒度下的分段效果。
原文摘要 · Abstract (English)
Dialogue topic segmentation supports summarization, retrieval, memory management, and conversational continuity. Despite decades of work, evaluation practice remains dominated by strict boundary matching and F1-based metrics. Modern large language model (LLM) based conversational systems increasingly rely on segmentation to manage conversation history beyond fixed context windows. In such systems, unstructured context accumulation degrades efficiency and coherence. This paper introduces an evaluation framework that reports boundary density and segment alignment diagnostics (purity and coverage) alongside window-tolerant F1 (W-F1). By separating boundary scoring from boundary selection, we evaluate segmentation quality across density regimes rather than at a single operating point. Cross-dataset evaluation shows that reported performance differences often reflect annotation granularity mismatch rather than boundary placement quality alone. We evaluate structurally distinct segmentation strategies across eight dialogue datasets spanning task-oriented, open-domain, meeting-style, and synthetic interactions. Boundary-based metrics are strongly coupled to boundary density: threshold sweeps produce larger W-F1 changes than switching between methods. These findings support viewing topic segmentation as a granularity selection problem rather than prediction of a single correct boundary set. This motivates separating boundary scoring from boundary selection for analyzing and tuning segmentation under varying annotation granularities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。