arXiv:2607.23442cs.CL2026-07

跨语言辩论中,中文模型更爱重复旧论点,英文则更善创新。

Do LLM Debates Repeat Arguments Differently Across Languages?

  • 用论点相似度衡量辩论中是否重复旧观点
  • 中文辩论的重复率显著高于英文,且在多模型下稳定存在
  • 适合关注多语言AI公平性与辩论质量评估的研究者

LLM辩论通常仅以最终答案评估,但对话记录可揭示后期是否发展新论点或以新措辞重提旧观点。本文通过‘先期论点相似度’分析,在71个议题、六种语言、四类模型的八轮辩论中发现:中文在三种多语言嵌入模型下,其先期论点相似度始终高于英文,且该差距在不同模型、轮次、回归调整、指标变体、提取长度控制、二次抽取子集及交叉编码尾部重评分等条件下均持续存在。人工校准显示:个体论点对齐弱,但高相似度尾部集中于实质性重复。引入多样性提示虽降低整体相似度,却未缩小中英文差距。因此,多语言辩论评估应追踪论点演进,并报告平均值与差距项。

原文摘要 · Abstract (English)

LLM debate is usually evaluated by final answers, yet transcripts reveal whether later turns develop new arguments or return to earlier claims in new wording. We study this process with \textit{prior-argument similarity}, which compares extracted argument units with earlier units in the same debate. In controlled eight-turn debates over 71 motions, six languages, and four model agents, Chinese is the only tested language with a consistently positive gap relative to English across three multilingual embedding models. The gap persists across agents, turn positions, regression adjustment, metric variants, extraction-length controls, a second-extractor subset, and cross-encoder tail rescoring. Manual calibration shows weak item-level alignment but a high-similarity tail enriched for substantive repetition. A diversity-aware prompt lowers \textit{prior-argument similarity} across languages, yet does not significantly narrow the Chinese--English gap. Multilingual debate evaluation should therefore measure argumentative development over time and report both average and gap terms.

大模型辩论跨语言比较论点重复多语言评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。