构建真实多轮混语对话数据集,推动多语言交流的自然语言处理研究
PingPong: A Natural Benchmark for Multi-Turn Code-Switching Dialogues
- 采集2-4人真实对话,涵盖五种语言组合与多线程结构
- 对话中回复常跨越多轮,平均回复距离更远,结构更复杂
- 提出问答、摘要、话题分类三项任务,验证现有模型性能不足
代码切换是全球多数双语或多语使用者日常交流中的普遍现象,但现有基准难以真实反映其复杂性。本文提出PingPong,一个面向多参与者、多轮次代码切换对话的自然语言基准,覆盖五种语言组合,部分为三语混合。数据集由人工撰写,包含2至4名参与者的真实对话,具有多线程结构,回复常追溯早期对话内容。实验表明,该数据在消息长度、说话人主导度和回复距离上比机器生成数据更具多样性与自然性。基于此,定义三项下游任务:问答、对话摘要与话题分类。对多个前沿语言模型的评估显示,其在代码切换输入上的表现仍有限,凸显开发更鲁棒的多语言处理系统刻不容缓。
原文摘要 · Abstract (English)
Code-switching is a widespread practice among the world's multilingual majority, yet few benchmarks accurately reflect its complexity in everyday communication. We present PingPong, a benchmark for natural multi-party code-switching dialogues covering five language-combination variations, some of which are trilingual. Our dataset consists of human-authored conversations among 2 to 4 participants covering authentic, multi-threaded structures where replies frequently reference much earlier points in the dialogue. We demonstrate that our data is significantly more natural and structurally diverse than machine-generated alternatives, offering greater variation in message length, speaker dominance, and reply distance. Based on these dialogues, we define three downstream tasks: Question Answering, Dialogue Summarization, and Topic Classification. Evaluations of several state-of-the-art language models on PingPong reveal that performance remains limited on code-switched inputs, underscoring the urgent need for more robust NLP systems capable of addressing the intricacies of real-world multilingual discourse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。