检验大模型摘要的可信度,发现不同模型生成结果差异大
Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers

- 设计双层诊断框架,评估摘要稳定性与语义一致性
- 三类文档上发现模型间生成差异显著,稳定性不一
- 适合教育场景中使用大模型摘要的研究者参考
利用大语言模型(LLM)进行零样本摘要已显著提升抽象摘要质量,但其内在随机性引发对生成摘要稳定性和可信度的担忧。随着学生和研究人员在教育场景中广泛采用零样本方式总结复杂学术材料,这一问题日益突出。本文提出一种两级诊断协议,用于评估LLM摘要器的可靠性:底层在受控环境下分析多份生成摘要的文档级稳定性,计算稳定性系数,并评估每份摘要与原文在语义和事实上的对齐程度;上层则通过分层抽样的文档样本,综合得出模型的整体稳定性指数,作为可信度代理指标。对三种LLM摘要器在三类文本上的实证研究显示,不同模型在生成层面的变异性存在统计显著差异。本研究首次以实证方式揭示了LLM摘要的稳定性问题,推动构建更稳健、可靠、可信的摘要系统。
原文摘要 · Abstract (English)
Zero-shot summarization using Large Language Models (LLMs) has significantly advanced the abstractive summarization task by producing coherent and fluent summaries. However, underlying stochasticity of the large language models raises concerns about the stability and trustworthiness of the LLM-generated summaries. This issue has become increasingly important due to proliferation of LLM-generated summaries in educational settings, where students and researchers summarize complex academic materials in zero-shot manner. We propose a novel two-level diagnostic protocol for benchmarking LLM-summarizers based on the stability of the generated summaries. At the lower level, document-level stability analysis is performed over multiple LLM-summaries generated under controlled environment, and the stability coefficient is computed. Each generated summary is scored for semantic and factual alignment with the original document, enabling estimation of stability along more than one dimensions. At the next level, observations from a stratified sample of documents drawn from the corpus are consolidated to estimate the stability index of the LLM-summarizer, which is the proxy for its trustworthiness. Our empirical investigation of three LLM-summarizers across three genres of documents reveals statistically significant differences in the generation-level variability among LLMs across summary evaluation metrics. This study advances the LLM-summarization research by evidential recognition of the stability problem in LLM-summaries and motivates further research towards development of robust, reliable and trustworthy LLM-summarizers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。