arXiv:2604.19689cs.AI2026-04被引 3

用分步推理计划提升艺术作品理解的可解释性与准确性

A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding

论文配图:A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding
图 1 · 摘自论文原文
  • 基于分步推理计划控制多模态检索,明确每步目标与证据需求
  • 在SemArt和Artpedia上优于传统检索与大模型基线,解释质量更高
  • 适合需要透明推理过程的文化类AI应用,如艺术教育与数字博物馆

理解艺术品需对视觉内容及文化、历史、风格背景进行多步推理。尽管近期多模态大语言模型在艺术解释方面展现潜力,但其依赖隐式推理与内化知识,限制了可解释性与显式证据支撑。我们提出A-MAR——一种基于智能体的多模态艺术检索框架,通过结构化推理计划显式引导检索过程。给定艺术品与用户查询,A-MAR首先将任务分解为包含目标与证据要求的结构化推理链;随后依据该计划进行检索,实现针对性证据选取,并支持逐步可追溯的解释。为评估艺术领域中的智能体式多模态推理,我们引入ArtCoT-QA诊断基准,包含多样艺术相关查询的多步推理链,可实现超越单一答案准确率的细粒度分析。在SemArt与Artpedia上的实验表明,A-MAR在最终解释质量上持续优于静态非规划检索与强基线多模态大模型;ArtCoT-QA评估进一步验证其在证据锚定与多步推理能力上的优势。结果凸显了推理条件化检索在知识密集型多模态理解中的重要性,推动构建可解释、目标驱动的AI系统,尤其适用于文化产业。代码与数据见:https://github.com/ShuaiWang97/A-MAR。

原文摘要 · Abstract (English)

Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large language models show promise in artwork explanation, they rely on implicit reasoning and internalized knowl- edge, limiting interpretability and explicit evidence grounding. We propose A-MAR, an Agent-based Multimodal Art Retrieval framework that explicitly conditions retrieval on structured reasoning plans. Given an artwork and a user query, A-MAR first decomposes the task into a structured reasoning plan that specifies the goals and evidence requirements for each step. Retrieval is then conditionedon this plan, enabling targeted evidence selection and supporting step-wise, grounded explanations. To evaluate agent-based multi- modal reasoning within the art domain, we introduce ArtCoT-QA. This diagnostic benchmark features multi-step reasoning chains for diverse art-related queries, enabling a granular analysis that extends beyond simple final answer accuracy. Experiments on SemArt and Artpedia show that A-MAR consistently outperforms static, non planned retrieval and strong MLLM baselines in final explanation quality, while evaluations on ArtCoT-QA further demonstrate its advantages in evidence grounding and multi-step reasoning ability. These results highlight the importance of reasoning-conditioned retrieval for knowledge-intensive multimodal understanding and position A-MAR as a step toward interpretable, goal-driven AI systems, with particular relevance to cultural industries. The code and data are available at: https://github.com/ShuaiWang97/A-MAR.

艺术理解多模态可解释性智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。