用AI生成档案图像描述,发现需人工审核才能保证质量。
ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing
- 用视觉语言模型生成考古图片目录,专家与普通人参与评估。
- AI生成内容准确度高时反而更难被识别,说明质量越好越需人工把关。
- 适合档案馆、博物馆人员参考,关注可信AI协作流程的设计。
摄影收藏品增长迅速,手动编目已难以跟上,促使使用视觉语言模型(VLMs)自动创建元数据。本研究考察了AI生成的目录描述能否接近人工撰写质量,并探讨生成式AI如何融入档案与博物馆的编目工作流。采用InternVL2模型为带有考古内容的印刷照片生成描述,由档案与考古专家及非专家在以人为中心的实验框架中进行评估。参与者需判断描述是AI生成还是专家撰写,评价质量,并表达对AI工具的使用意愿与信任度。分类表现显著高于随机水平,但两组均低估了自身识别能力。OCR错误与幻觉降低了感知质量,而准确度和实用性更高的描述反而更难被识别,表明对开箱即用模型生成内容仍需人工审查,尤其在考古等专业领域。专家采纳意愿较低,主要担忧在于保存责任而非技术性能。研究主张采用协作模式:AI负责初稿生成,人类负责验证,确保符合策展价值(如来源真实性、透明性)。成功整合不仅依赖领域微调等技术进步,更取决于专业人士的信任建立,可通过透明可解释的AI流程共同促进。
原文摘要 · Abstract (English)
The accelerating growth of photographic collections has outpaced manual cataloguing, motivating the use of vision language models (VLMs) to automate metadata generation. This study examines whether Al-generated catalogue descriptions can approximate human-written quality and how generative Al might integrate into cataloguing workflows in archival and museum collections. A VLM (InternVL2) generated catalogue descriptions for photographic prints on labelled cardboard mounts with archaeological content, evaluated by archive and archaeology experts and non-experts in a human-centered, experimental framework. Participants classified descriptions as AI-generated or expert-written, rated quality, and reported willingness to use and trust in AI tools. Classification performance was above chance level, with both groups underestimating their ability to detect Al-generated descriptions. OCR errors and hallucinations limited perceived quality, yet descriptions rated higher in accuracy and usefulness were harder to classify, suggesting that human review is necessary to ensure the accuracy and quality of catalogue descriptions generated by the out-of-the-box model, particularly in specialized domains like archaeological cataloguing. Experts showed lower willingness to adopt AI tools, emphasizing concerns on preservation responsibility over technical performance. These findings advocate for a collaborative approach where AI supports draft generation but remains subordinate to human verification, ensuring alignment with curatorial values (e.g., provenance, transparency). The successful integration of this approach depends not only on technical advancements, such as domain-specific fine-tuning, but even more on establishing trust among professionals, which could both be fostered through a transparent and explainable AI pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。