首个面向繁体中文的多模态评测基准,评估模型在图文音频任务中的表现与速度。
Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan
- 构建900道繁体中文多模态题目,覆盖图像、音频与文本组合
- 闭源模型整体优于开源模型,但开源模型在音频任务中表现不俗
- 端到端架构比分步处理更省时,适合低延迟应用
多模态大语言模型(MLLMs)能处理视觉、听觉和文本输入,克服单模态大模型的局限。然而现有评测基准常忽略繁体中文下的三模态评估,且未考虑推理延迟。为此,我们提出Multi-TW,首个针对繁体中文的多模态模型性能与延迟评测基准。Multi-TW包含900道多选题,数据源自由台湾语能力测验指导委员会(SC-TOP)设计的官方水平测试,涵盖图像+文本、音频+文本配对。我们评估了多种任意模态间(any-to-any)模型及带音频转录的视觉语言模型(VLMs)。结果表明,闭源模型在各类模态中普遍优于开源模型,但开源模型在音频任务上表现良好。端到端的任意模态流水线相比使用独立音频转录的VLM具有显著延迟优势。Multi-TW全面呈现模型能力,凸显繁体中文微调与高效多模态架构的必要性。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) process visual, acoustic, and textual inputs, addressing the limitations of single-modality LLMs. However, existing benchmarks often overlook tri-modal evaluation in Traditional Chinese and do not consider inference latency. To address this, we introduce Multi-TW, the first Traditional Chinese benchmark for evaluating the performance and latency of any-to-any multimodal models. Multi-TW includes 900 multiple-choice questions (image and text, audio and text pairs) sourced from official proficiency tests developed with the Steering Committee for the Test of Proficiency-Huayu (SC-TOP). We evaluated various any-to-any models and vision-language models (VLMs) with audio transcription. Our results show that closed-source models generally outperform open-source ones across modalities, although open-source models can perform well in audio tasks. End-to-end any-to-any pipelines offer clear latency advantages compared to VLMs using separate audio transcription. Multi-TW presents a comprehensive view of model capabilities and highlights the need for Traditional Chinese fine-tuning and efficient multimodal architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。