提出新指标衡量科学新闻对读者的知识提升效果
KnowledgeGain: Evaluating and Optimizing Science News Generation for Reader Learning

- 用读者知识增长量评估新闻质量,而非仅看语义相似度
- 模拟人类阅读者筛选文章,使读后准确率提升12.7%
- 适合关注科普传播效果与认知目标的研究者
科学新闻是科研成果向公众传播的重要媒介。现有文本生成或摘要评价多关注语义相似性和事实一致性,但未能衡量读者实际获得的知识量。本文提出KnowledgeGain指标,通过测量读者阅读后的知识增量来评估科学新闻质量。通过控制实验验证该指标能有效捕捉不同媒体形式带来的知识差异,并据此校准出仅依赖提示的大型语言模型阅读者模拟器。该模拟器用于候选文章的排序与过滤,减少人工评估负担。第二次人类实验显示,经模拟器筛选的文章在读后准确性及标准化KnowledgeGain上均优于强基线模型,提升12.7%。本工作推动科学新闻生成向符合布卢姆分类学认知目标的方向迈进。
原文摘要 · Abstract (English)
Science news is an important medium to communicate discoveries between the research communities and the public. Yet, most metrics for generated or summarized text evaluate semantic similarity and factual consistency, but do not measure how much knowledge readers learn from the news. We introduce KnowledgeGain, a metric that evaluates the quality of science news by measuring how much knowledge readers gained after reading it. To evaluate the metric, we first performed a controlled human study and showed that the metric successfully captures the differential knowledge gained by human readers reading different types of science media. The data allowed us to calibrate a prompt-only LLM reader simulator. We use it to rank and filter candidate articles before human evaluation. A second human study shows that articles selected with this simulator improve post-reading accuracy and normalized KnowledgeGain over a strong generation baseline. Our work is a step toward generating science news that better meets the knowledge and comprehension goals of Bloom's Taxonomy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。