arXiv:2506.00332cs.CLcs.SI2025-06ACL被引 3

首个标注的多语言混用聊天语料库,助力自然对话研究。

Disentangling Codemixing in Chats: The NUS ABC Codemixed Corpus

  • 构建含35万+条消息的标注语料库,支持多语言混用分析。
  • 涵盖英、中文等多语言模式,覆盖真实聊天场景。
  • 适合计算语言学、社会语言学与NLP应用研究者使用。

代码混用是指在单一话语中无缝融合多种语言元素,反映自然的多语言交流模式。尽管在社交媒体、聊天消息和即时通讯中广泛存在,但缺乏公开可用、经作者标注且适合建模人类对话关系的语料库。本研究推出了首个标注的通用语料库,用于理解上下文中的代码混用,同时严格遵守隐私与伦理标准。该项目持续收集、验证并整合代码混用消息,以JSON格式发布,附带详细元数据与语言统计信息。目前语料库已包含超过355,641条消息,涵盖英语、普通话及其他语言的多种混用模式。该语料库有望成为计算语言学、社会语言学及NLP应用研究的基础数据集。

原文摘要 · Abstract (English)

Code-mixing involves the seamless integration of linguistic elements from multiple languages within a single discourse, reflecting natural multilingual communication patterns. Despite its prominence in informal interactions such as social media, chat messages and instant-messaging exchanges, there has been a lack of publicly available corpora that are author-labeled and suitable for modeling human conversations and relationships. This study introduces the first labeled and general-purpose corpus for understanding code-mixing in context while maintaining rigorous privacy and ethical standards. Our live project will continuously gather, verify, and integrate code-mixed messages into a structured dataset released in JSON format, accompanied by detailed metadata and linguistic statistics. To date, it includes over 355,641 messages spanning various code-mixing patterns, with a primary focus on English, Mandarin, and other languages. We expect the Codemix Corpus to serve as a foundational dataset for research in computational linguistics, sociolinguistics, and NLP applications.

语料库多语言代码混用NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。