用人类思维导图提升零样本多模态推理,防作弊、懂上下文。
CogGuide: Human-Like Guidance for Zero-Shot Omni-Modal Reasoning
- 构建‘理解-规划-选择’三模块流程,用意图草图引导推理
- 零微调下跨模型提升最高达9.51个百分点,抑制错误捷径
- 适合需要可靠零样本推理的多模态应用开发者
针对多模态大模型在复杂跨模态推理中存在‘捷径’和上下文理解不足的问题,本文提出一种基于人类认知策略的零样本多模态推理组件——CogGuide,核心为‘意图草图’。该组件采用即插即用的三模块流水线:意图感知器、策略生成器、策略选择器,显式构建‘理解-规划-选择’的认知过程。通过生成并筛选‘意图草图’策略来引导最终推理,无需参数微调,仅通过上下文工程实现跨模型迁移。信息论分析表明,该过程可降低条件熵,提升信息利用效率,从而抑制非预期的捷径推理。在IntentBench、WorldSense和Daily-Omni上的实验验证了方法的通用性与鲁棒增益;相比各自基线,完整三模块方案在不同推理引擎与流水线组合中均表现一致提升,最高达约9.51个百分点,证明了‘意图草图’推理组件在零样本场景下的实用价值与可移植性。
原文摘要 · Abstract (English)
Targeting the issues of "shortcuts" and insufficient contextual understanding in complex cross-modal reasoning of multimodal large models, this paper proposes a zero-shot multimodal reasoning component guided by human-like cognitive strategies centered on an "intent sketch". The component comprises a plug-and-play three-module pipeline-Intent Perceiver, Strategy Generator, and Strategy Selector-that explicitly constructs a "understand-plan-select" cognitive process. By generating and filtering "intent sketch" strategies to guide the final reasoning, it requires no parameter fine-tuning and achieves cross-model transfer solely through in-context engineering. Information-theoretic analysis shows that this process can reduce conditional entropy and improve information utilization efficiency, thereby suppressing unintended shortcut reasoning. Experiments on IntentBench, WorldSense, and Daily-Omni validate the method's generality and robust gains; compared with their respective baselines, the complete "three-module" scheme yields consistent improvements across different reasoning engines and pipeline combinations, with gains up to approximately 9.51 percentage points, demonstrating the practical value and portability of the "intent sketch" reasoning component in zero-shot scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。