用双重控制机制提升大模型深度研究的准确性和可追溯性
RhinoInsight: Improving Deep Research through Control Mechanisms for Model Behavior and Context
- 将用户需求拆解为可验证的分步目标,防止无效规划
- 通过证据审计减少幻觉,提升内容可信度
- 适合需要高可靠性的科研写作与复杂决策场景
大语言模型正从单轮应答转向具备持续推理与决策能力的工具型智能体。现有系统采用计划-搜索-写作-报告的线性流程,因缺乏对模型行为和上下文的显式控制,易产生错误累积与上下文衰减。我们提出RhinoInsight框架,引入两种控制机制以增强鲁棒性、可追溯性与整体质量,且无需参数更新。首先,可验证清单模块将用户需求转化为可追踪、可验证的子目标,结合人工或大模型评审进行优化,并生成层级化提纲,锚定后续行动,避免不可执行的规划。其次,证据审计模块结构化搜索内容,迭代更新提纲并剔除噪声上下文,同时由评审者对高质量证据排序并绑定至撰写内容,确保可验证性并降低幻觉。实验表明,RhinoInsight在深度研究任务中达到当前最优表现,同时在深度搜索任务上保持竞争力。
原文摘要 · Abstract (English)
Large language models are evolving from single-turn responders into tool-using agents capable of sustained reasoning and decision-making for deep research. Prevailing systems adopt a linear pipeline of plan to search to write to a report, which suffers from error accumulation and context rot due to the lack of explicit control over both model behavior and context. We introduce RhinoInsight, a deep research framework that adds two control mechanisms to enhance robustness, traceability, and overall quality without parameter updates. First, a Verifiable Checklist module transforms user requirements into traceable and verifiable sub-goals, incorporates human or LLM critics for refinement, and compiles a hierarchical outline to anchor subsequent actions and prevent non-executable planning. Second, an Evidence Audit module structures search content, iteratively updates the outline, and prunes noisy context, while a critic ranks and binds high-quality evidence to drafted content to ensure verifiability and reduce hallucinations. Our experiments demonstrate that RhinoInsight achieves state-of-the-art performance on deep research tasks while remaining competitive on deep search tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。