arXiv:2510.12740cs.CLcs.AI2025-10Conference of the …被引 1

用对话敏感性评估语言模型自然度,发现其更倾向延续核心话题。

Hey, wait a minute: on at-issue sensitivity in Language Models

  • 将对话分段生成续写,重组后比较概率判断自然度。
  • 模型更偏好延续当前讨论焦点,尤其在指令微调后更明显。
  • 适合研究对话理解与生成的学者,关注语用敏感性建模。

评估语言模型对话自然度颇具挑战:自然度标准不一,且缺乏可扩展的量化指标。本文借助语言学中的‘在议性’(at-issueness)概念,提出新方法 Divide, Generate, Recombine, and Compare(DGRC)。该方法将对话分割为片段作为提示,用语言模型生成各部分的后续内容,再重组并比较重组序列的似然值。此方法降低语言模型分析中的偏差,支持对话语敏感行为的系统测试。实验表明,语言模型倾向于延续‘在议性’内容,且指令微调模型表现更显著;当出现如“嘿,等一下”等提示时,这种偏好会减弱。尽管指令微调未进一步增强调节能力,但该模式体现了有效对话的核心特征。

原文摘要 · Abstract (English)

Evaluating the naturalness of dialogue in language models (LMs) is not trivial: notions of 'naturalness' vary, and scalable quantitative metrics remain limited. This study leverages the linguistic notion of 'at-issueness' to assess dialogue naturalness and introduces a new method: Divide, Generate, Recombine, and Compare (DGRC). DGRC (i) divides a dialogue as a prompt, (ii) generates continuations for subparts using LMs, (iii) recombines the dialogue and continuations, and (iv) compares the likelihoods of the recombined sequences. This approach mitigates bias in linguistic analyses of LMs and enables systematic testing of discourse-sensitive behavior. Applying DGRC, we find that LMs prefer to continue dialogue on at-issue content, with this effect enhanced in instruct-tuned models. They also reduce their at-issue preference when relevant cues (e.g., "Hey, wait a minute") are present. Although instruct-tuning does not further amplify this modulation, the pattern reflects a hallmark of successful dialogue dynamics.

对话生成语言模型语用分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。