首个支持多方言对话生成的框架,让AI真正听懂各地英语口音。
DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English
- 构建九种英语方言的平行对话数据集,涵盖词汇、拼写、语法三要素。
- 人类评估显示98.8%偏好该数据,模型对非标准英语识别率不足70%。
- 适合研究方言识别、跨文化对话系统及大模型后训练的开发者。
超过80%的16亿英语使用者不使用标准美式英语(SAE),但大语言模型常无法识别非SAE方言,生成刻板回应。本文提出DialectLLM,首个大规模方言感知对话生成框架,覆盖书面方言三大要素:词汇、拼写与形态句法。该框架生成涵盖九种英语方言的方言平行对话数据集。通过与母语语言学家合作,设计并验证了从SAE到各方言的转换规则,确保真实性。挑战了当前将单一形态句法特征应用于用户输入与模型输出的惯例,发现模型不应复制高达90%的方言语法特征。人工评估确认数据质量,标注者在98.8%的对比中更偏好DialectLLM数据。进一步构建了包含5万+对话、9.7万+问答对的DialectLLM-Bench基准,评估17个大模型在方言识别与生成任务中的表现。即使前沿模型准确率也低于70%,加拿大英语等显著方言识别率不足50%,且系统性地将非美英方言误判为美式或英式。此外,实验证明该数据可作为大模型后训练的有效资源,为实现方言感知对话系统提供可行路径。
原文摘要 · Abstract (English)
More than 80% of the 1.6B English speakers do not use Standard American English (SAE), yet LLMs often fail to correctly identify non-SAE dialects and generate stereotyped responses for their speakers. We introduce DialectLLM, the first large-scale framework for generating high-quality multi-dialectal conversational data encompassing the three pillars of written dialect -- lexical (vocabulary), orthographic (spelling), and morphosyntactic (grammar) features. DialectLLM produces a dialect-parallel dialog dataset spanning nine English dialects. Partnering with native linguists, we design and validate SAE-to-dialect transformation rules, ensuring authenticity. Our approach challenges the prevailing practice of applying a single morphosyntactic feature set to both user utterances and model responses, showing that models should not reproduce up to 90% of the grammatical features of a dialect. Human evaluation confirms data quality, with annotators preferring DialectLLM over prior methods in 98.8% of pairwise comparisons for dialect naturalness. We then construct DialectLLM-Bench, a dialect-parallel benchmark with 50k+ dialogs, resulting in 97k+ QA pairs, and evaluate 17 LLMs on dialect identification and response generation tasks. Even frontier models achieve under 70% accuracy, fail to reach 50% for prominent dialects like Canadian English, and systematically misclassify non-SAE dialects as American or British. Beyond benchmarking, we show that DialectLLM data also serve as a scalable LLM post-training resource, suggesting a practical path toward dialect-aware conversational AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。