arXiv:2602.19458cs.AIcs.HC2026-02被引 1

用互补信息训练大模型,帮决策者发现独特有用的新信号。

ComplLLM: Fine-tuning LLMs to Discover Complementary Signals for Decision-making

  • 基于决策理论,用互补信息作为奖励微调大模型
  • 在真实和模拟任务中成功恢复已知互补信号
  • 适合需要多视角决策支持的专家系统

多智能体决策流程在各智能体具备互补性时,表现优于单一智能体。我们提出 ComplLLM,一种基于决策理论的后训练框架,通过将互补信息作为奖励,微调决策辅助大模型,使其输出能补充现有智能体决策的独特信号。我们在涉及领域专家的真实与合成任务中验证了该方法,结果表明其能有效恢复已知的互补信息,并生成合理可信的解释,为下游决策者提供支持。

原文摘要 · Abstract (English)

Multi-agent decision pipelines can outperform single agent workflows when complementarity holds, i.e., different agents bring unique information to the table to inform a final decision. We propose ComplLLM, a post-training framework based on decision theory that fine-tunes a decision-assistant LLM using complementary information as reward to output signals that complement existing agent decisions. We validate ComplLLM on synthetic and real-world tasks involving domain experts, demonstrating how the approach recovers known complementary information and produces plausible explanations of complementary signals to support downstream decision-makers.

大模型决策支持多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。