arXiv:2605.26620cs.CLcs.HC2026-05中稿 · EMNLP被引 1

提出无需参考的文本粒度衡量方法,可分析问答系统表现差异。

Granuscore: A Reference-Free Measure of Granularity for Text Analysis and Question Answering

论文配图:Granuscore: A Reference-Free Measure of Granularity for Text Analysis and Question Answering
图 1 · 摘自论文原文
  • 利用嵌入空间层级结构定义无参考粒度评分
  • 在多个数据集上准确捕捉语义粒度差异
  • 适合评估问答模型输出与数据集难度

自然语言表达的信息粒度各异,从细粒度指称到宏观描述。现有度量多关注表面细节或句子特定性。本文提出 Granuscore,一种基于层次嵌入空间结构特性的无参考粒度度量方法。该方法在 Granola-EQ 数据集上可靠恢复了层级顺序,并捕捉到不同语篇情境下的预期粒度差异。跨领域分析显示,Granuscore 能解释句子长度之外的非线性句式特定性变化。进一步应用于四个问答基准测试,分析问题、标准答案与模型输出在不同响应结果下的粒度差异,揭示模型行为的一致性模式,为刻画问答数据集难度提供理论依据。整体表明,Granuscore 是一种可扩展、广泛适用的文本粒度分析工具。

原文摘要 · Abstract (English)

Natural language conveys information at varying levels of granularity, from fine-grained references to broad descriptions. While granularity is fundamental to human communication, existing measures mostly capture surface detail or sentence specificity. We introduce Granuscore, a reference-free measure of granularity that leverages structural properties of a hierarchical embedding space. Granuscore reliably recovers hierarchical orderings on the Granola-EQ dataset and captures expected differences in granularity across discourse contexts. Across domains, we further show that Granuscore explains non-linear variation in sentence specificity beyond sentence length. Finally, we apply Granuscore to four question-answering benchmarks and analyze how granularity differs for questions, gold answers, and model outputs across response outcomes. The analysis reveals consistent differences in model behavior and provides a principled lens for characterizing the difficulty of QA datasets. Together, the results position Granuscore as a scalable, broadly applicable tool for analyzing granularity in text.

文本粒度问答系统嵌入分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。