arXiv:2603.27043cs.CL2026-03中稿 · LREC 2026

开源双语访谈数据集,含中英对话与语言态度分析。

Introducing MELI: the Mandarin-English Language Interview Corpus

  • 收集51名双语者中英对照录音,涵盖朗读与自由访谈。
  • 总时长约29.8小时,中英文各约14.7与15.1小时,支持跨语言对比。
  • 适合语音研究、语言态度与代码转换分析的学者使用。

我们介绍MELI语料库,一个开放资源,包含51名普通话-英语双语者的29.8小时语音数据。该语料库结合了普通话和英语的匹配会话,涵盖朗读句子与关于语言变体、标准性及学习经历的自发访谈。音频采样率为44.1 kHz(16位,立体声)。所有访谈均完成转录,并在词与音素层面进行强制对齐,且已匿名化处理。描述性统计显示,普通话部分总计约14.7小时(平均时长17.3分钟),英语部分约15.1小时(平均时长17.8分钟)。报告了每种语言的词元/词类统计量,并记录了代码转换模式(普通话会话中频繁出现,英语会话中较少)。语料库设计支持跨说话人、跨语言的声学比较,并将声学特征与说话人陈述的语言态度关联,支持定量与定性分析。MELI语料库将连同转录文本、对齐信息、元数据、标注地图扫描件及文档,在CC BY-NC 4.0许可下发布。

原文摘要 · Abstract (English)

We introduce the Mandarin-English Language Interview (MELI) Corpus, an open-source resource of 29.8 hours of speech from 51 Mandarin-English bilingual speakers. MELI combines matched sessions in Mandarin and English with two speaking styles: read sentences and spontaneous interviews about language varieties, standardness, and learning experiences. Audio was recorded at 44.1 kHz (16-bit, stereo). Interviews were fully transcribed, force-aligned at word and phone levels, and anonymized. Descriptively, the Mandarin component totals ~14.7 hours (mean duration 17.3 minutes) and the English component ~15.1 hours (mean duration 17.8 minutes). We report token/type statistics for each language and document code-switching patterns (frequent in Mandarin sessions; more limited in English sessions). The corpus design supports within-/cross-speaker, within/cross-language acoustic comparison and links acoustics to speakers' stated language attitudes, enabling both quantitative and qualitative analyses. The MELI Corpus will be released with transcriptions, alignments, metadata, scans of labelled maps and documentation under a CC BY-NC 4.0 license.

语音语料双语研究语言态度代码转换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。