测试大模型仅靠内部知识生成书摘的准确性与一致性。
Evaluating book summaries from internal knowledge in Large Language Models: a cross-model and semantic consistency approach
- 用多个大模型生成书摘,再让其他模型互评,避免偏见。
- 结果表明不同模型在内容和风格上差异明显,有的更贴近人类总结。
- 适合研究大模型知识表征与自动评估方法的人参考。
我们研究大语言模型(LLMs)仅依靠内部知识生成完整准确的书摘能力,不依赖原文。采用多样书籍和多种模型架构,考察模型能否合成与人类已有理解一致的有意义叙事。通过LLM-as-a-judge范式进行评估:每个模型生成的摘要与高质量人工摘要对比,所有参与模型同时评价自己及他人输出。该方法可识别潜在偏差,如模型偏好自身风格。使用ROUGE和BERTScore量化人类摘要与模型摘要之间的对齐程度,评估语法和语义对应深度。结果显示各模型在内容呈现和风格偏好上存在细微差异,揭示了依赖内部知识进行摘要任务时的优缺点。研究有助于深入理解大模型对事实信息的内部编码机制及跨模型评估动态,对构建更稳健的自然语言生成系统具有意义。
原文摘要 · Abstract (English)
We study the ability of large language models (LLMs) to generate comprehensive and accurate book summaries solely from their internal knowledge, without recourse to the original text. Employing a diverse set of books and multiple LLM architectures, we examine whether these models can synthesize meaningful narratives that align with established human interpretations. Evaluation is performed with a LLM-as-a-judge paradigm: each AI-generated summary is compared against a high-quality, human-written summary via a cross-model assessment, where all participating LLMs evaluate not only their own outputs but also those produced by others. This methodology enables the identification of potential biases, such as the proclivity for models to favor their own summarization style over others. In addition, alignment between the human-crafted and LLM-generated summaries is quantified using ROUGE and BERTScore metrics, assessing the depth of grammatical and semantic correspondence. The results reveal nuanced variations in content representation and stylistic preferences among the models, highlighting both strengths and limitations inherent in relying on internal knowledge for summarization tasks. These findings contribute to a deeper understanding of LLM internal encodings of factual information and the dynamics of cross-model evaluation, with implications for the development of more robust natural language generative systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。