arXiv:2505.13559cs.CLcs.LG2025-05被引 2

首个多语言混码对话摘要基准,揭示大模型理解混码的深层缺陷

CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models

  • 构建跨语种混码对话摘要基准CS-Sum,覆盖中英、泰英、马来英三对语言
  • 尽管自动指标分数高,但大模型仍频繁出现改变对话意义的细微错误
  • 发现三类典型错误,提示需针对性训练以提升混码理解能力

混码现象对大语言模型构成重大挑战,但其可理解性在现有研究中尚未充分探索。本文提出CS-Sum,通过将混码对话转译为英文摘要来评估大模型对混码的理解能力。CS-Sum是首个针对中文-英语(EN-ZH)、泰米尔语-英语(EN-TA)和马来语-英语(EN-MS)三组语言对的混码对话摘要基准,每对语言包含900至1300条人工标注对话。我们评估了十种大模型(含开源与闭源),涵盖少样本、翻译-摘要和微调(基于合成数据的LoRA、QLoRA)等方法。结果表明,尽管自动化指标得分较高,但大模型仍存在影响语义的细微错误。本文归纳出三类最常见的错误类型,且错误率在不同语言对与模型间差异显著,凸显了针对混码数据进行专门训练的必要性。

原文摘要 · Abstract (English)

Code-switching (CS) poses a significant challenge for Large Language Models (LLMs), yet its comprehensibility remains underexplored in LLMs. We introduce CS-Sum, to evaluate the comprehensibility of CS by the LLMs through CS dialogue to English summarization. CS-Sum is the first benchmark for CS dialogue summarization across Mandarin-English (EN-ZH), Tamil-English (EN-TA), and Malay-English (EN-MS), with 900-1300 human-annotated dialogues per language pair. Evaluating ten LLMs, including open and closed-source models, we analyze performance across few-shot, translate-summarize, and fine-tuning (LoRA, QLoRA on synthetic data) approaches. Our findings show that though the scores on automated metrics are high, LLMs make subtle mistakes that alter the complete meaning of the dialogue. To this end, we introduce 3 most common type of errors that LLMs make when handling CS input. Error rates vary across CS pairs and LLMs, with some LLMs showing more frequent errors on certain language pairs, underscoring the need for specialized training on code-switched data.

混码对话大模型评估摘要生成语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。