用视觉语言模型自动提取拍卖目录中的物品信息,助力艺术史研究。
Lot Machine: Multimodal Lot Extraction from Auction Catalogs

- 基于视觉语言模型,设计多策略提示与解码框架提取目录信息。
- 商用模型表现最优,本地部署需强制输出结构以保证格式正确。
- 适合文化遗产机构在隐私和资源限制下进行大规模数据挖掘。
为艺术史溯源与市场研究,拍卖目录是追踪特定物品时空轨迹的重要资源。尽管历史拍卖目录遵循领域惯例,其内部排版却高度不一致,且缺乏机器可读的条目级元数据,制约了大规模分析。本文提出一个端到端管道,从19至20世纪德国拍卖销售目录(German Sales)中自动提取结构化物品级元数据。基于人工标注的代表性页面测试集,评估了多种提示策略与约束解码框架下的视觉语言模型(VLMs)表现。考虑文化机构的实际约束(预算、算力、数据隐私),在从商业接口到本地量化模型的不同部署模式下进行基准测试。结果表明,商用接口达到性能上限,而机构网关提供可行且保护隐私的替代方案;本地部署虽可行,但必须在生成时强制输出结构以确保返回有效JSON。尽管仍需不同程度的人工校正,本工作证明基于VLM的流程可成功实现历史拍卖目录的大规模自动化分析。
原文摘要 · Abstract (English)
For provenance research and art market studies, auction catalogs are an essential resource to trace specific objects over time and space. While historical auction catalogs follow established domain conventions, their internal formatting remains highly variable, and their large-scale analysis is currently restricted by the lack of machine-readable representations of the auction lots. We propose a pipeline to automatically extract structured lot-level metadata from German Sales, a large database of historical auction and sales catalogs from the 19th and 20th centuries. Using a manually annotated test set of representative catalog pages, we evaluate Vision-Language Models (VLMs) under varying prompt strategies and constrained decoding frameworks. To reflect the practical constraints faced by cultural heritage institutions, including budget, compute resources, and data privacy requirements, we benchmark the methods across different deployment modes ranging from commercial providers to locally hosted, quantized models. We find that commercial endpoints establish the performance ceiling, while institutional gateways offer a viable, privacy-preserving alternative. Local deployments remain feasible, but strictly require enforcing the output structure during generation to guarantee a valid JSON format. While varying degrees of human-in-the-loop correction are still necessary, this work demonstrates that a VLM-based pipeline can successfully unlock historical auction catalogs for large-scale automated analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。