arXiv:2603.00105cs.LGcs.CL2026-03

用分层视角评估大模型摘要质量,可解释关键词更可信

LIDS: LLM Summary Inference Under the Layered Lens

  • 基于BERT-SVD方向度量,量化摘要与原文相似性
  • 通过重复提示捕捉统计不确定性,提升评估鲁棒性
  • 可识别关键主题词,适合模型对比与可解释性研究

自2022年ChatGPT问世以来,大语言模型在自然语言处理领域受到广泛关注。其生成摘要的能力尤为突出,但摘要质量评估仍因语言复杂性而困难。本文提出一种新的摘要评估方法LIDS,结合BERT-SVD方向度量与SOFARI(LIDS),可衡量摘要准确性并提供分层主题的可解释关键词。LIDS利用BERT嵌入和重复提示,通过潜在SVD方向度量计算摘要与原文相似性,并量化统计不确定性,实现大文本摘要的自然嵌入表示。进一步运用SOFARI方法,在控制错误发现率(FDR)的前提下,挖掘每个潜在主题对应的关键词。大量实证研究通过人工验证和与其他相似性度量方法的对比,证明了LIDS在实际应用中的有效性与鲁棒性,涵盖不同大模型的比较。

原文摘要 · Abstract (English)

Large language models (LLMs) have gained significant attention by many researchers and practitioners in natural language processing (NLP) since the introduction of ChatGPT in 2022. One notable feature of ChatGPT is its ability to generate summaries based on prompts. Yet evaluating the quality of these summaries remains challenging due to the complexity of language. To this end, in this paper we suggest a new method of LLM summary inference with BERT-SVD-based direction metric and SOFARI (LIDS) that assesses the summary accuracy equipped with interpretable key words for layered themes. The LIDS uses a latent SVD-based direction metric to measure the similarity between the summaries and original text, leveraging the BERT embeddings and repeated prompts to quantify the statistical uncertainty. As a result, LIDS gives a natural embedding of each summary for large text reduction. We further exploit SOFARI to uncover important key words associated with each latent theme in the summary with controlled false discovery rate (FDR). Comprehensive empirical studies demonstrate the practical utility and robustness of LIDS through human verification and comparisons to other similarity metrics, including a comparison of different LLMs.

大模型评估可解释性摘要质量BERT-SVD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。