arXiv:2503.10427cs.CL2025-03被引 3

首个面向台湾繁体中文的视觉语言模型评测基准

VisTW: Benchmarking Vision-Language Models for Traditional Chinese in Taiwan

  • 构建双模块评测集:多选题测试知识推理,对话题检验文化语境理解
  • 涵盖21个学科、131组图文对,覆盖台湾本土文化语境
  • 揭示主流模型在繁体中文视觉任务中的显著性能差距

本文提出首个针对台湾繁体中文的视觉语言模型(VLM)综合评测基准。该评估体系包含两个互补组件:(1) VisTW-MCQ,由21个学科的手动精选多选题组成,用于测试VLM的广泛知识与推理能力;(2) VisTW-Dialogue,包含131组手动创建的图像-问题对,用于评估VLM在台湾文化语境下的自由对话生成能力。现有评测多集中于英文或简体中文,忽视了台湾、香港等地繁体中文的独特语言与文化特征。我们的分析揭示了不同VLM在处理繁体中文视觉内容时的显著性能差异,并指出了具体挑战。

原文摘要 · Abstract (English)

In this paper, we propose a comprehensive evaluation benchmark for Visual Language Models (VLM) in Traditional Chinese. Our evaluation suite, the first of its kind, contains two complementary components: (1) VisTW-MCQ, a collection of manually curated exam multi-choice questions from 21 academic subjects designed to test the broad knowledge and reasoning capabilities of VLMs; and (2) VisTW-Dialogue, an open dialogue benchmark comprising 131 image-question pairs manually created to evaluate VLMs' ability in free-form dialogue generation within Taiwanese cultural contexts. These benchmarks address a critical gap in the evaluation landscape, where existing benchmarks predominantly focus on English or Simplified Chinese, neglecting the unique linguistic and cultural aspects of Traditional Chinese used in regions like Taiwan and Hong Kong. Our analysis reveals significant performance differences across various VLMs and highlights specific challenges in processing Traditional Chinese visual content.

视觉语言模型繁体中文评测基准台湾文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。