让AI从回答出发,生成新见解,提升开放问答体验。
An Answer is just the Start: Related Insight Generation for Open-Ended Document-Grounded QA

- 用聚类构建文档主题图谱,再通过邻域选择找相关上下文。
- 在3000个问题上验证,生成的见解既相关又多样且可操作。
- 适合研究开放问答、人机交互与信息探索的学者和开发者。
开放问答对AI而言仍具挑战,因其需超越事实检索,涉及综合判断与探索,用户常通过多轮迭代优化答案。现有评测基准未支持此过程。为此,我们提出新任务:基于文档集合的相关见解生成,目标是从文档中提炼出能改进、拓展或重构初始答案的新见解,从而增强用户交互与问答体验。我们构建并发布了SCOpE-QA(科学文献开放问答数据集),包含20个研究领域的3000个开放问题。提出InsightGen方法,分两阶段:先用聚类生成文档主题表示,再基于主题图谱的邻域选择获取相关上下文,利用大模型生成多样化且相关的见解。在3000个问题上,使用两种生成模型与两种评估设置进行广泛测试,结果表明InsightGen持续产出有用、相关且可操作的见解,为该新任务建立了坚实基线。
原文摘要 · Abstract (English)
Answering open-ended questions remains challenging for AI systems because it requires synthesis, judgment, and exploration beyond factual retrieval, and users often refine answers through multiple iterations rather than accepting a single response. Existing QA benchmarks do not explicitly support this refinement process. To address this gap, we introduce a new task, document-grounded related insight generation, where the goal is to generate additional insights from a document collection that help improve, extend, or rethink an initial answer to an open-ended question, ultimately supporting richer user interaction and a better overall question answering experience. We curate and release SCOpE-QA (Scientific Collections for Open-Ended QA), a dataset of 3,000 open-ended questions across 20 research collections. We present InsightGen, a two-stage approach that first constructs a thematic representation of the document collection using clustering, and then selects related context based on neighborhood selection from the thematic graph to generate diverse and relevant insights using LLMs. Extensive evaluation on 3,000 questions using two generation models and two evaluation settings shows that InsightGen consistently produces useful, relevant, and actionable insights, establishing a strong baseline for this new task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。