重构关键点分析的结构,提升总结质量与覆盖率。
Key Point Analysis Needs Structure Recovery: Task Definition, Dataset Diagnosis, and a Structure-Aware Benchmark

- 将关键点分析定义为结构化预测任务,需恢复语义分组。
- 新基准覆盖更全,关键点更准确,流行度估计更可靠。
- 适合研究可解释性、评估方法及大模型评判的学者。
关键点分析(KPA)旨在识别一组能概括多条论点及其出现频率的精炼要点。我们指出,KPA本质上是需要恢复语义分组、生成代表性要点、确保覆盖全面并准确估算流行度的结构化预测问题。现有基准在分组质量、冗余、覆盖度和论点-要点映射上存在缺陷,导致基于参考的评估出现天花板效应与选择失败。为此,我们通过人机协同重标注构建了一个结构感知、分布敏感的新基准。人类与大模型评估均显示,新标注结构在分组连贯性、要点质量、覆盖度和流行度估计方面显著优于现有数据。我们还发布了多个标注资源,支持论点-要点匹配、可解释性KPA、LLM作为裁判等研究,并提出真正KPA的研究路线图。
原文摘要 · Abstract (English)
Key Point Analysis (KPA) aims to identify a concise set of key points that summarize a collection of arguments together with their prevalence. We argue that KPA is fundamentally a structured prediction problem that requires recovering semantic groupings, generating representative key points, ensuring coverage, and estimating prevalence. Under this formulation, we show that existing KPA benchmarks suffer from limitations in grouping quality, redundancy, coverage, and argument-key point mappings, causing ceiling violation and selection failure in reference-based evaluation. To support future research on true KPA, we introduce a structure-aware, distribution-sensitive benchmark built via a human-in-the-loop re-annotation. Human and LLM evaluations consistently show that the resulting structures yield more coherent groupings, higher-quality key points, better coverage, and more reliable prevalence estimates than existing annotations. We further release several annotation resources to support research on KPA evaluation, argument-key point matching, explainable KPA, and LLM-as-a-judge methodologies, and outline a research agenda for true KPA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。