arXiv:2509.25946cs.AI2025-09NeurIPS被引 3

用多模态多步推理自动发现既保细节又泛化的最佳模型。

Automated Model Discovery via Multi-modal & Multi-step Pipeline

  • 通过双视觉语言模块分步分析并提出候选模型
  • 在细节捕捉与整体泛化上均优于现有方法
  • 适合需要自动化模型选择的研究者和工程师

自动化模型发现是指在大规模组合搜索空间中,自动搜索并识别适用于给定数据集的最佳模型。现有方法往往难以在捕捉细微特征与保证训练外泛化能力之间取得平衡,同时控制合理模型复杂度。本文提出一种多模态、多步推理的自动化模型发现流程。该流程利用两个视觉-语言模型(VLM)——AnalyzerVLM 和 EvaluatorVLM——以代理方式实现模型提议与评估。AnalyzerVLM 自主规划并执行多步分析,提出有效的候选模型;EvaluatorVLM 则从定量与感知两方面评估模型在局部细节适应性与整体趋势泛化性上的表现。实验表明,该流程能有效发现兼具细节捕捉能力和强泛化性的模型。大量消融研究进一步验证了多模态与多步推理在发现优质模型中的关键作用。

原文摘要 · Abstract (English)

Automated model discovery is the process of automatically searching and identifying the most appropriate model for a given dataset over a large combinatorial search space. Existing approaches, however, often face challenges in balancing the capture of fine-grained details with ensuring generalizability beyond training data regimes with a reasonable model complexity. In this paper, we present a multi-modal \& multi-step pipeline for effective automated model discovery. Our approach leverages two vision-language-based modules (VLM), AnalyzerVLM and EvaluatorVLM, for effective model proposal and evaluation in an agentic way. AnalyzerVLM autonomously plans and executes multi-step analyses to propose effective candidate models. EvaluatorVLM assesses the candidate models both quantitatively and perceptually, regarding the fitness for local details and the generalibility for overall trends. Our results demonstrate that our pipeline effectively discovers models that capture fine details and ensure strong generalizability. Additionally, extensive ablation studies show that both multi-modality and multi-step reasoning play crucial roles in discovering favorable models.

自动化建模多模态智能代理模型搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。