arXiv:2603.28474cs.CVcs.AI2026-03

用AI分析古瓷器,能识朝代、窑口、纹样等六项特征,还支持解释性描述。

CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains

  • 结合视觉与多模态检索,通过工具增强推理分析瓷器细节。
  • 在6个属性上平均比GPT-5高12.2%准确率,超越所有开源闭源模型。
  • 专为瓷器鉴赏设计,适合文博研究者和文化爱好者使用。

古瓷器鉴赏需要深厚的历史知识、材料理解与审美感知,非专业者难以参与。为促进文化遗产普及并辅助专家鉴赏,我们提出CiQi-Agent——面向古瓷器鉴赏的领域专用智能代理。该模型支持多图输入,可调用视觉与多模态检索工具,实现对朝代、年号、窑口、釉色、纹饰、器型六项属性的细粒度分析。它不仅能捕捉细微视觉特征,还能检索领域知识,融合视觉与文本证据生成连贯可解释的鉴赏描述。为此,我们构建了大规模专家标注数据集CiQi-VQA,包含29,596件瓷器、51,553张图像及557,940个视觉问答对,并建立对应基准CiQi-Bench。CiQi-Agent通过监督微调、强化学习与工具增强推理框架训练。实验表明,其7B版本在CiQi-Bench上六项属性均优于各类开源与闭源模型,平均准确率高出GPT-5 12.2%。模型与数据已公开发布于https://huggingface.co/datasets/SII-Monument-Valley/CiQi-VQA。

原文摘要 · Abstract (English)

The connoisseurship of antique Chinese porcelain demands extensive historical expertise, material understanding, and aesthetic sensitivity, making it difficult for non-specialists to engage. To democratize cultural-heritage understanding and assist expert connoisseurship, we introduce CiQi-Agent -- a domain-specific Porcelain Connoisseurship Agent for intelligent analysis of antique Chinese porcelain. CiQi-Agent supports multi-image porcelain inputs and enables vision tool invocation and multimodal retrieval-augmented generation, performing fine-grained connoisseurship analysis across six attributes: dynasty, reign period, kiln site, glaze color, decorative motif, and vessel shape. Beyond attribute classification, it captures subtle visual details, retrieves relevant domain knowledge, and integrates visual and textual evidence to produce coherent, explainable connoisseurship descriptions. To achieve this capability, we construct a large-scale, expert-annotated dataset CiQi-VQA, comprising 29,596 porcelain specimens, 51,553 images, and 557,940 visual question--answering pairs, and further establish a comprehensive benchmark CiQi-Bench aligned with the previously mentioned six attributes. CiQi-Agent is trained through supervised fine-tuning, reinforcement learning, and a tool-augmented reasoning framework that integrates two categories of tools: a vision tool and multimodal retrieval tools. Experimental results show that CiQi-Agent (7B) outperforms all competitive open- and closed-source models across all six attributes on CiQi-Bench, achieving on average 12.2\% higher accuracy than GPT-5. The model and dataset have been released and are publicly available at https://huggingface.co/datasets/SII-Monument-Valley/CiQi-VQA.

古瓷器多模态智能鉴赏领域代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。