arXiv:2509.16438cs.CVcs.CL2025-09中稿 · ArabicNLP 2025

用大模型自动构建阿拉伯语视频文本检索基准,省去大量人工校对。

AutoArabic: A Three-Stage Framework for Localizing Video-Text Retrieval Benchmarks

  • 三阶段框架用大模型翻译非阿拉伯语基准,减少近四倍人工修改。
  • 生成40,144条流畅阿拉伯语描述,错误检测准确率达97%。
  • 适用于多语言基准本地化研究,代码开源可复现。

视频到文本和文本到视频的检索任务主要依赖英文基准(如DiDeMo、MSR-VTT)及近期多语言语料库(如RUDDER),但阿拉伯语仍缺乏本地化评估指标。本文提出三阶段框架AutoArabic,利用先进大语言模型(LLMs)将非阿拉伯语基准自动翻译为现代标准阿拉伯语,使人工修订量减少近四倍。框架包含一个错误检测模块,可自动标记潜在翻译错误,准确率达97%。将该框架应用于DiDeMo视频检索基准,生成了包含40,144条流畅阿拉伯语描述的DiDeMo-AR。通过分析翻译错误并建立分类体系,为未来阿拉伯语本地化提供指导。在阿拉伯语与英语版本上使用相同超参数训练CLIP风格基线模型,发现两者在Recall@1上仅存在约3个百分点的性能差距,表明阿拉伯语本地化保留了基准难度。评估三种后编辑预算(零/仅标记项/完全编辑)发现,性能随编辑量增加而提升,且原始大模型输出(零预算)仍具可用性。为确保其他语言可复现,代码已公开于https://github.com/Tahaalshatiri/AutoArabic。

原文摘要 · Abstract (English)

Video-to-text and text-to-video retrieval are dominated by English benchmarks (e.g. DiDeMo, MSR-VTT) and recent multilingual corpora (e.g. RUDDER), yet Arabic remains underserved, lacking localized evaluation metrics. We introduce a three-stage framework, AutoArabic, utilizing state-of-the-art large language models (LLMs) to translate non-Arabic benchmarks into Modern Standard Arabic, reducing the manual revision required by nearly fourfold. The framework incorporates an error detection module that automatically flags potential translation errors with 97% accuracy. Applying the framework to DiDeMo, a video retrieval benchmark produces DiDeMo-AR, an Arabic variant with 40,144 fluent Arabic descriptions. An analysis of the translation errors is provided and organized into an insightful taxonomy to guide future Arabic localization efforts. We train a CLIP-style baseline with identical hyperparameters on the Arabic and English variants of the benchmark, finding a moderate performance gap (about 3 percentage points at Recall@1), indicating that Arabic localization preserves benchmark difficulty. We evaluate three post-editing budgets (zero/ flagged-only/ full) and find that performance improves monotonically with more post-editing, while the raw LLM output (zero-budget) remains usable. To ensure reproducibility to other languages, we made the code available at https://github.com/Tahaalshatiri/AutoArabic.

视频检索多语言自动化阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。