构建多尺度小说摘要幻觉检测基准,助力长文本生成可靠性研究
LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization

- 基于29部中文小说与BookSum章节数据,构建中英双语多尺度幻觉检测数据集
- 设计8类幻觉类型,通过多模型仲裁与实体引导生成确保数据平衡与真实
- 手动修订测试集内容,提升可靠性,适合长文本生成与幻觉研究者使用
尽管近年来上下文窗口显著扩展,长文本摘要中的幻觉问题仍难解决。长篇小说因事件细节丰富、对话详尽,比新闻或论文更适合作为幻觉研究对象。然而现有研究缺乏针对长文本小说摘要的多尺度幻觉检测基准,也未充分探索幻觉随上下文增长的变化规律。本文提出LongNovel,一个涵盖29部中文小说(16k至100k token)及BookSum章节级数据的多尺度、中英双语小说幻觉检测基准。设计8类幻觉类型,采用多模型仲裁与实体参考幻觉生成方法,确保数据真实性与类别分布均衡。同时对测试集进行人工修正以保障可靠性。大量实验表明LongNovel具有挑战性,已开源供后续研究使用。
原文摘要 · Abstract (English)
Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues. However, current research lacks a multi-scale benchmark for hallucination detection in long-context novel summarization and does not fully explore how hallucinations change as the context grows longer. In this study, we propose LongNovel, a multi-scale long-context bilingual (Chinese and English) novel benchmark for hallucination detection. This benchmark is constructed from 29 Chinese novels (ranging from 16k to 100k tokens) and chapter-level data from the BookSum dataset. We design 8 hallucination types and employ a combination of Multi-Model Arbitration and Entity-Referenced Hallucination Generation to ensure both data authenticity and a balanced distribution of hallucination categories. Furthermore, we manually revise the content in the test set to guarantee data reliability. Extensive experimental results demonstrate that LongNovel is a challenging benchmark. We release LongNovel for future research. https://github.com/BDML-lab/LongNovel
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。