arXiv:2603.27435cs.CLcs.AI2026-03被引 3

让大模型理解写作意图,生成更高质量的长篇问答报告。

Improving Attributed Long-form Question Answering with Intent Awareness

  • 用标签化结构提取隐含写作意图,引导模型生成
  • 大模型零样本生成提升2.9分,小模型提升12.3分
  • 适合需要高可读性与精准引用的科研报告场景

大型语言模型(LLMs)被广泛用于生成知识密集型的综合性报告。然而,这些模型虽训练于多样学术论文与报告,却未接触作者撰写文档时的推理过程与写作意图。我们假设增强模型的意图感知能力可显著提升生成报告的质量。为此,我们设计并采用结构化的标签方案,更有效地激发和提取写作或引用背后的隐含意图。实验表明,这些提取的意图不仅能提升大模型在零样本情况下的生成能力,还能用于生成高质量合成数据以微调小型模型。在多个具有挑战性的科学报告生成任务中,大模型平均性能提升2.9个百分点,小模型提升12.3个百分点。此外,分析显示,意图感知能优化模型引用使用方式,并大幅提高报告可读性。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being used to generate comprehensive, knowledge-intensive reports. However, while these models are trained on diverse academic papers and reports, they are not exposed to the reasoning processes and intents that guide authors in crafting these documents. We hypothesize that enhancing a model's intent awareness can significantly improve the quality of generated long-form reports. We develop and employ structured, tag-based schemes to better elicit underlying implicit intents to write or cite. We demonstrate that these extracted intents enhance both zero-shot generation capabilities in LLMs and enable the creation of high-quality synthetic data for fine-tuning smaller models. Our experiments reveal improved performance across various challenging scientific report generation tasks, with an average improvement of +2.9 and +12.3 absolute points for large and small models over baselines, respectively. Furthermore, our analysis illuminates how intent awareness enhances model citation usage and substantially improves report readability.

长篇问答意图建模生成质量LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。