用可验证证据链破解艺术隐性影响难题,避免盲目匹配。
M-ArtAgent: Evidence-Based Multimodal Agent for Implicit Art Influence Discovery
- 构建四阶段推理框架,结合图像与传记证据链
- 在WIB-100数据集上达成83.7%的F1值和0.910的ROC-AUC
- 引入历史约束与对抗验证,适合艺术史与AI交叉研究者
隐性艺术影响虽视觉上合理,但常无文字记载,导致归属困难:相似性是必要条件而非充分证据。以往系统多依赖嵌入相似性或标签驱动图补全,而近期多模态大模型仍存在时间不一致与未验证归属问题。本文提出M-ArtAgent,一种基于证据的多模态智能体,将隐性影响发现重构为概率性裁决。其采用四阶段协议——调查、佐证、反驳与判决,由类ReAct控制器统筹,从图像与传记中整合可验证证据链,强制遵循艺术史公理,并通过提示隔离的批评者对每项假设进行对抗性反驳。两个理论驱动算子——StyleComparator(用于沃尔夫林形式分析)与ConceptRetriever(用于ICONCLASS图标学基础)——确保中间结论可形式化审计。在包含100位艺术家与2,000个有向关系的平衡版WikiArt Influence Benchmark-100(WIB-100)上,M-ArtAgent实现83.7%正类F1、0.666 Matthews相关系数(MCC)与0.910的受试者工作特征曲线下面积(ROC-AUC),且在屏蔽显式影响表述后,性能仍保持稳定。结果表明,隐性影响分析需历史根基裁决,而非仅靠模式匹配。
原文摘要 · Abstract (English)
Implicit artistic influence, although visually plausible, is often undocumented and thus poses a historically constrained attribution problem: resemblance is necessary but not sufficient evidence. Most prior systems reduce influence discovery to embedding similarity or label-driven graph completion, while recent multimodal large language models (LLMs) remain vulnerable to temporal inconsistency and unverified attributions. This paper introduces M-ArtAgent, an evidence-based multimodal agent that reframes implicit influence discovery as probabilistic adjudication. It follows a four-phase protocol consisting of Investigation, Corroboration, Falsification, and Verdict governed by a Reasoning and Acting (ReAct)-style controller that assembles verifiable evidence chains from images and biographies, enforces art-historical axioms, and subjects each hypothesis to adversarial falsification via a prompt-isolated critic. Two theory-grounded operators, StyleComparator for Wolfflin formal analysis and ConceptRetriever for ICONCLASS-based iconographic grounding, ensure that intermediate claims are formally auditable. On the balanced WikiArt Influence Benchmark-100 (WIB-100) of 100 artists and 2,000 directed pairs, M-ArtAgent achieves 83.7% positive-class F1, 0.666 Matthews correlation coefficient (MCC), and 0.910 area under the receiver operating characteristic curve (ROC-AUC), with leakage-control and robustness checks confirming that the gains persist when explicit influence phrases are masked. By coupling multimodal perception with domain-constrained falsification, M-ArtAgent demonstrates that implicit influence analysis benefits from historically grounded adjudication rather than pattern matching alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。