构建首个大规模印地语系多语言混合对话数据集
IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
- 用多语言大模型生成带人物设定的多轮对话
- 涵盖18种方言,超132万条自然混码对话
- 适合研究印地语系低资源语言对话系统
大型语言模型(LLMs)已重塑对话AI,但高质量多语言混合对话资源仍匮乏,尤其在印地语系语言中,使用者常在本地文字和罗马化形式间自然切换。我们提出IndicTalk,目前最大规模的印地语系多语言混合对话语料库,包含18种语言变体、9种印地语系语言,覆盖超过132.86万条基于事件的多轮对话。该语料库通过全自动流程生成:结合真实新闻事件作为上下文,利用多语言大模型进行人物设定引导的对话生成,并自动完成质量验证。大量语言学分析、自动评估与人工评测表明,IndicTalk生成的对话流畅、连贯且自然混合多种语言。我们将公开该数据集,以支持低资源印地语系语言对话AI的发展与评估。数据集地址:https://huggingface.co/datasets/LingoIITGN/IndicTalk。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。