arXiv:2510.14949cs.CLcs.CV2025-10被引 4

测试6种英语方言对生成模型的影响,发现普通方法效果差,新方法能显著提升方言表现。

DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation

论文配图:DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
图 1 · 摘自论文原文
  • 用编码器设计新策略,让模型识别方言特征同时不损失标准英语表现。
  • 在5种方言上性能提升34.4%,接近标准英语水平,对标准语几乎无影响。
  • 适用于希望提升方言兼容性的图像视频生成系统开发者。

英语等接触语言存在丰富的地区变体(方言),方言使用者常以方言与生成模型互动。当前多模态生成模型能否有效处理方言输入?本文构建涵盖六种常见英语方言的大规模基准数据集,通过方言使用者收集并验证超过4200个独特提示,评估17个图像与视频生成模型。自动与人工评估结果显示,仅使用一个方言词导致模型性能下降32.26%至48.17%。现有缓解方法如微调和提示重写仅能带来小于7%的改善,且可能严重损害标准美式英语(SAE)表现。为此,我们提出一种基于编码器的通用缓解策略,使模型在学习识别新方言特征的同时保持SAE性能。在Stable Diffusion 1.5等模型上的实验表明,该方法可使五种方言性能达到与SAE相当水平(+34.4%),对SAE性能影响近乎为零。

原文摘要 · Abstract (English)

Contact languages like English exhibit rich regional variations in the form of dialects, which are often used by dialect speakers interacting with generative models. However, can multimodal generative models effectively produce content given dialectal textual input? In this work, we study this question by constructing a new large-scale benchmark spanning six common English dialects. We work with dialect speakers to collect and verify over 4200 unique prompts and evaluate on 17 image and video generative models. Our automatic and human evaluation results show that current state-of-the-art multimodal generative models exhibit 32.26% to 48.17% performance degradation when a single dialect word is used in the prompt. Common mitigation methods such as fine-tuning and prompt rewriting can only improve dialect performance by small margins (< 7%), while potentially incurring significant performance degradation in Standard American English (SAE). To this end, we design a general encoder-based mitigation strategy for multimodal generative models. Our method teaches the model to recognize new dialect features while preserving SAE performance. Experiments on models such as Stable Diffusion 1.5 show that our method is able to simultaneously raise performance on five dialects to be on par with SAE (+34.4%), while incurring near zero cost to SAE performance.

多模态生成方言鲁棒性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。