提出无需参考的文本粒度衡量方法,可分析问答系统表现差异。
Granuscore: A Reference-Free Measure of Granularity for Text Analysis and Question Answering

- 利用嵌入空间层级结构定义无参考粒度评分
- 在多个数据集上准确捕捉语义粒度差异
- 适合评估问答模型输出与数据集难度
自然语言表达的信息粒度各异,从细粒度指称到宏观描述。现有度量多关注表面细节或句子特定性。本文提出 Granuscore,一种基于层次嵌入空间结构特性的无参考粒度度量方法。该方法在 Granola-EQ 数据集上可靠恢复了层级顺序,并捕捉到不同语篇情境下的预期粒度差异。跨领域分析显示,Granuscore 能解释句子长度之外的非线性句式特定性变化。进一步应用于四个问答基准测试,分析问题、标准答案与模型输出在不同响应结果下的粒度差异,揭示模型行为的一致性模式,为刻画问答数据集难度提供理论依据。整体表明,Granuscore 是一种可扩展、广泛适用的文本粒度分析工具。
原文摘要 · Abstract (English)
Natural language conveys information at varying levels of granularity, from fine-grained references to broad descriptions. While granularity is fundamental to human communication, existing measures mostly capture surface detail or sentence specificity. We introduce Granuscore, a reference-free measure of granularity that leverages structural properties of a hierarchical embedding space. Granuscore reliably recovers hierarchical orderings on the Granola-EQ dataset and captures expected differences in granularity across discourse contexts. Across domains, we further show that Granuscore explains non-linear variation in sentence specificity beyond sentence length. Finally, we apply Granuscore to four question-answering benchmarks and analyze how granularity differs for questions, gold answers, and model outputs across response outcomes. The analysis reveals consistent differences in model behavior and provides a principled lens for characterizing the difficulty of QA datasets. Together, the results position Granuscore as a scalable, broadly applicable tool for analyzing granularity in text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。