通过优化因果链与互信息,提升模型抽象推理能力
DIO: Refining Mutual Information and Causal Chain to Enhance Machine Abstract Reasoning Ability
- 构建因果链模型,从图像到答案分步建模推理过程
- 在RPM测试中准确率显著提升,首次实现开放问答生成
- 适合关注抽象推理与可解释性研究的学者
尽管深度学习广泛应用,其抽象推理能力仍存瓶颈。本文以瑞文渐进矩阵(RPM)为基准,建模完整的因果链:图像 → 属性 → 进展模式 → 一致性 → 答案,并提出基线模型DIO。然而,DIO采用的互信息下界优化目标存在局限:边界松散且基于统计,忽视了因果主客体关联。为此,本文提出三项改进:1)引入可训练负样本(Brando),收紧变分下界;2)用高斯混合特征模型替代生成,提供无限加权负样本(WORLD),进一步紧化边界;3)加入元数据监督(DIEGO),弥合属性到模式间的语义鸿沟,使表征更符合人类规则。改进后模型在判别式RPM任务中表现大幅提升,并首次实现开放题型的有效答案生成。研究提供了因果驱动的设计范式、目标优化策略及跨模态洞察。
原文摘要 · Abstract (English)
Despite deep learning's broad success, its abstract-reasoning bottleneck persists. We tackle Raven's Progressive Matrices (RPM), the benchmark for pattern, reasoning and problem-solving intelligence. We model the full causal chain image $\rightarrow$ attributes $\rightarrow$ progressive patterns $\rightarrow$ consistency $\rightarrow$ answer and build the baseline DIO. Yet DIO's mutual-information lower-bound objective does not embed human logic: the bound is loose and statistic-based, ignoring causal subject-object links. We therefore present three refinements. 1) Brando introduces trainable negative options to tighten the variational bound. 2) WORLD replaces generation with a Gaussian-mixture feature model that supplies infinite, weighted negatives, further tightening the bound. 3) DIEGO adds metadata supervision to rectify the "attributes $\rightarrow$ patterns" semantic gap, aligning representations with human rules. These upgrades substantially boost discriminative RPM accuracy and, for the first time, let DIO generate valid answers in open-ended RPM. The work provides causal-driven design guidelines, objective-refinement strategies and cross-modal insights for abstract-reasoning research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。