arXiv:2605.31080cs.MMcs.AI2026-05

用小模型为视障者生成多语言艺术描述,效果因语言而异。

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models

论文配图:A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
图 1 · 摘自论文原文
  • 用小规模视觉语言模型+语言特化适配器生成多语言艺术描述。
  • 德语下多语言适配表现最好,罗、塞语则单语言适配更稳定。
  • 适合关注视障者跨语言艺术可及性的研究与应用落地者。

视障者在博物馆等场景中仍缺乏多语言艺术描述支持,尤其受限于隐私和知识产权时,更倾向于使用本地部署的小型视觉语言模型(VLM)。本试点研究采用 Qwen2.5-VL-3B-Instruct 模型,在德语、罗马尼亚语和塞尔维亚语上探索策展人引导的多语言艺术描述。我们构建了基于作品图像与元数据的视障友好平行描述语料库,对比语言特化 LoRA 适配器与单一多语言适配器在固定主干和训练预算下的表现。评估结合自动词法与嵌入指标,以及经少量罗马尼亚语视障用户试点校准的 LLM-as-Judge 协议。结果显示:在罗、塞语中,语言特化适配器在可控性与视觉相关性上更稳定;德语中多语言适配表现相当。这些发现为小型本地化 VLM 的部署提供实证依据,强调需更大规模视障用户研究与更广语言覆盖,方可推广结论。

原文摘要 · Abstract (English)

Blind and low-vision (BLV) audiences remain underserved by visual art descriptions, particularly across languages and in museum settings where privacy and intellectual-property constraints may favour small on-premise vision-language models (VLMs). This pilot study investigates curator-guided multilingual art description with Qwen2.5-VL-3B-Instruct for German, Romanian, and Serbian. We construct a parallel BLV-oriented caption corpus from artwork images and metadata, and compare language-specific LoRA adapters with a single multilingual adapter under a fixed backbone and training budget. Evaluation combines automatic lexical and embedding-based metrics with an LLM-as-Judge protocol calibrated against a small Romanian BLV pilot study. Under our pilot setup, language-specific adapters show more stable controllability and visually grounded description quality for Romanian and Serbian, while multilingual adaptation remains competitive in German. We frame these findings as deployment-oriented evidence for small on-premise VLMs, and highlight the need for larger BLV user studies and broader language coverage before drawing general conclusions about multilingual accessibility.

视障辅助多语言小模型艺术描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。