arXiv:2605.16364cs.SDcs.AI2026-05被引 1

构建首个阿拉伯语真实场景语音交互数据集,含用户反馈与答案可回答性标注。

WASIL: In-the-Wild Arabic Spoken Interactions with LLMs

论文配图:WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
图 1 · 摘自论文原文
  • 基于多ASR一致性的后编辑生成低成本高质量转录文本
  • 收录8529条语音交互,14.2%为负面反馈,覆盖四种方言和标准语
  • 可区分识别错误与固有不可回答问题,适合语音助手优化研究

大型语言模型(LLMs)语音助手通常采用自动语音识别(ASR)到LLM的级联架构,识别错误会扭曲用户意图。不满意的反馈也可能源于模糊、域外或非请求性对话轮次,难以分离出ASR的影响。我们发布了WASIL(阿拉伯语中意为连接或关联):包含音频、ASR候选、助手回复及明确喜欢/不喜欢反馈的真实场景阿拉伯语语音交互数据集(8,529个对话轮次;14.2%为不喜欢),并提供一个2,000轮测试集,涵盖现代标准阿拉伯语(MSA)及四种主要方言及其标签。通过多ASR一致性引导的后编辑生成低成本高精度转录文本,并标注回答可回答性(可回答、模糊/需澄清、不支持、非请求/噪声),以区分内在不可回答性与由ASR引起的性能下降。最后,提出一种基于多裁判大模型评分的无参考式响应评估方法,用于比较来自ASR转录与原始语音的响应质量。

原文摘要 · Abstract (English)

Large Language Models (LLMs) voice assistants are commonly built as cascaded Automatic Speech recognition (ASR) to LLM systems, where recognition errors can distort user intent. Dislikes may also arise from ambiguous, out-of-domain, or non-request turns, making it hard to isolate ASR effects. We release WASIL (it denotes connection or linking in Arabic): in-the-wild Arabic spoken interaction prompts with audio, ASR hypotheses, assistant responses, and explicit like/dislike feedback (8,529 turns; 14.2% dislikes), plus a 2,000-turn test set covering Modern Standard Arabic (MSA) and four major dialects with their labels. We provide low-cost gold transcripts via multi-ASR agreement-guided post-editing and annotate answerability (answerable, ambiguous/needs-clarification, unsupported, not-a-request/noise) to separate intrinsic unanswerability from ASR-induced degradation. Finally, we describe scalable reference-free evaluation of responses from ASR vs. gold transcripts using multi-judge LLM scoring.

语音交互阿拉伯语大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。