提出新评估框架,精准区分科学叙事中的创意与幻觉。
Hallucination or Creativity: How to Evaluate AI-Generated Scientific Stories?
- 构建融合语义对齐与实体幻觉检测的综合评分体系。
- 发现现有方法难以区分教学式重构与事实错误。
- 适合评估AI生成科普故事的质量与可信度。
生成式AI可将科学论文转化为面向多样受众的叙述性内容,但评估这些故事仍具挑战。讲故事需要抽象、简化和教学创意,而这些特性常无法被标准摘要指标捕捉。同时,科学语境下事实幻觉至关重要,但现有检测器常误判合法的叙述重构,或在涉及创意时表现不稳定。本文提出StoryScore,一个整合语义对齐、词汇锚定、叙事控制、结构保真、冗余规避及实体级幻觉检测的复合评估框架。分析揭示,许多幻觉检测方法失效的根本原因在于:尽管自动指标能有效评估与原文的语义相似性,却难以衡量其叙述方式与控制能力。
原文摘要 · Abstract (English)
Generative AI can turn scientific articles into narratives for diverse audiences, but evaluating these stories remains challenging. Storytelling demands abstraction, simplification, and pedagogical creativity-qualities that are not often well-captured by standard summarization metrics. Meanwhile, factual hallucinations are critical in scientific contexts, yet, detectors often misclassify legitimate narrative reformulations or prove unstable when creativity is involved. In this work, we propose StoryScore, a composite metric for evaluating AI-generated scientific stories. StoryScore integrates semantic alignment, lexical grounding, narrative control, structural fidelity, redundancy avoidance, and entity-level hallucination detection into a unified framework. Our analysis also reveals why many hallucination detection methods fail to distinguish pedagogical creativity from factual errors, highlighting a key limitation: while automatic metrics can effectively assess semantic similarity with original content, they struggle to evaluate how it is narrated and controlled.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。