用专家知识与可视化结合,让艺术史家高效构建古画语义图谱。
VisTCP: A Visualization Framework to Construct Knowledge-Graph-Based Representation for Traditional Chinese Painting

- 融合领域专家知识与智能模型,生成古画语义结构化表示。
- 通过可视化对比专家标注与模型预测,支持人工修正与迭代优化。
- 适合艺术史研究者、文物数字化团队使用,提升古画理解可信度。
结构化表示能刻画图像中的语义对象及其关系,为传统中国绘画(TCP)的语义理解提供有效途径,助力考古与艺术史研究。然而,现有面向图像的结构化表示方法在处理TCP时表现不佳,主要受两大挑战制约:一是TCP中的对象与事件与现代自然图像差异显著,易引发语义误解;二是古代对象与事件的精准识别对领域专家也极具难度。本文提出VisTCP,一种结合面向TCP的智能模型与专家知识的可视化框架,使艺术史家可在人机协同模式下构建可信的结构化表示。首先,我们联合三位领域专家开展预研,建立TCP语义分类体系;随后,利用专家标注数据训练面向TCP的结构化表示模型,可自动提取绘画中的有意义对象及其关系。为体现模型不确定性,设计联合嵌入可视化视图,呈现专家标注与模型预测的差异,帮助用户基于领域知识修正结果,实现模型的迭代优化。最后,通过真实数据集上的案例研究、使用场景及专家访谈,验证了VisTCP在支持TCP结构化表示与语义理解方面的有效性。
原文摘要 · Abstract (English)
Structured representation can characterize semantic objects and relationships in images. It provides a possible effective way for the semantic understanding of Traditional Chinese Paintings (TCPs) to better support archaeology and art history research. However, most image-oriented structured representation methods perform poorly on TCPs, due to two major challenges: 1) the objects and events of TCPs exhibit substantial differences from modern natural images, which results in semantic misunderstandings of TCPs; and 2) it is difficult to achieve accurate identification of ancient objects and events in TCPs, even for domain experts.In this paper, we propose VisTCP, a visualization framework that combines a TCP-oriented intelligent model and expert knowledge, which enables art historians to achieve trustworthy structured representations of TCPs in a human-in-the-loop manner. Firstly, we conduct a pilot study with three domain experts to build a semantic taxonomy of TCPs. Then, expert-annotated data are used to train a TCP-oriented structured representation model, which can automatically extract meaningful objects and their relationships in TCPs. To inform users of the model uncertainty, we design a joint embedding visualization view to show the differences between expert annotations and model predictions. This allows users to refine the structured representation based on their domain knowledge, enabling iterative optimization of the model. Finally, we conduct a case study, a usage scenario, and expert interviews on a real dataset to demonstrate the effectiveness of VisTCP in supporting the structured representation and semantic understanding of TCPs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。