arXiv:2604.19245cs.CLcs.AI2026-04中稿 · ACL被引 1

研究大模型在对话修复中的不可靠行为,揭示其多轮交互的不稳定性

Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs

论文配图:Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
图 1 · 摘自论文原文
  • 通过数学题对话测试模型自我修复与响应用户修复的能力
  • 不同模型对修复请求反应差异显著,部分易被误导,部分完全抗拒
  • 多轮对话后模型行为更难预测,每种模型有独特不可靠模式

修复是人类对话中解决误解的重要机制,但在人-大模型交互中仍被忽视。本文研究大模型在可解与不可解数学问题的多轮对话中如何参与修复过程,考察模型是否主动发起修复及如何回应用户发起的修复。结果显示,各模型表现差异明显:从几乎完全抗拒修复到极易被引导,行为呈现显著不一致性。进一步发现,当对话超过单轮后,模型行为变得更加独特且难以预测。总体表明,每个测试的大模型在修复情境下均表现出独特的不可靠性特征。

原文摘要 · Abstract (English)

Repair, an important resource for resolving trouble in human-human conversation, remains underexplored in human-LLM interaction. In this study, we investigate how LLMs engage in the interactive process of repair in multi-turn dialogues around solvable and unsolvable math questions. We examine whether models initiate repair themselves and how they respond to user-initiated repair. Our results show strong differences across models: reactions range from being almost completely resistant to (appropriate) repair attempts to being highly susceptible and easily manipulated. We further demonstrate that once conversations extend beyond a single turn, model behavior becomes more distinctive and less predictable across systems. Overall, our findings indicate that each tested LLM exhibits its own characteristic form of unreliability in the context of repair.

大模型行为对话修复不可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。