让科学答案更易懂:为关键决策设计可沟通的检索系统
Relevance is not enough: A Communication-Oriented Retrieval System for Consequential Scientific Question Answering

- 按推理类型分类问题,生成追问以补全证据
- 用通俗语言重写科学内容,降低理解门槛
- 强调实际影响和情感因素,适合非专家使用
AI系统越来越多地回答健康、安全与环境等科学问题。但现有检索增强生成系统多关注事实准确性和话题相关性,忽视了非专业人士对答案意义的理解需求。本文聚焦直接影响人们生活的‘重要性科学问题’,通过一个公众水质沟通系统研究发现:仅答对问题仍不足,需解释与背景,且风险信息的情感影响不可忽略。系统首先按推理类型(如因果或政策类)分类问题,再生成追问以识别缺失证据并检索。一模块明确区分已知与未知,另一模块将科学内容转化为易懂语言,采用‘关心邻居’或‘行政官员’等角色化语气调节表达风格。在超过160个问题上的消融实验表明,该系统使用约十分之一的上下文信息,且在多种配置下提升了人类评估的完整性。与社区成员共同设计的完整性指标及微调后的学习评判器显示,标准相关性得分仅解释约1%的人类完整性评分差异,即使优化后也仅中度吻合,说明完整性是当前指标无法可靠捕捉的人类中心目标。
原文摘要 · Abstract (English)
AI systems increasingly answer scientific questions about health, safety, and the environment. But most retrieval-augmented generation systems are tuned to provide factually correct, on-topic answers rather than to help non-experts understand what those answers mean for their lives and decisions. We focus on consequential scientific questions whose results directly shape people's lives and study them through a public water-quality communication system, where residents and community leaders interpret the findings and choose actions. Their experiences show that on-topic answers can still be insufficient without explanation and context and that the emotional weight of risk information cannot be ignored. Our system first classifies each question by reasoning type (for example, causal versus policy-based), then generates follow-up questions to identify missing evidence and retrieve it. One component clearly distinguishes between what is known and what is uncertain, while another rewrites scientific details into accessible language, using persona-based styles, such as a caring neighbor or an administrative official, to adapt tone and readability. Ablations on over $160$ questions show that the system uses $\textit{an order of magnitude less context}$ and, in several configurations, improves human-rated completeness. A completeness metric co-designed with community members and a fine-tuned learned judge reveal that standard relevance scores explain about $1\%$ of variation in human completeness ratings, and even the tuned judge only moderately aligns with humans, indicating that completeness is a distinct human-centered objective that current metrics do not reliably capture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。