arXiv:2506.21497cs.CL2025-06被引 1

用用户后续反应指导对话模型,提升社交对话中的参与度。

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments

  • 通过用户互动后的反馈信号,直接优化对话模型。
  • 在情感支持和说服对话中,用户参与度显著提升。
  • 适合研究人机交互与对话系统优化的学者。

在社交驱动型对话中,通过互动提升用户参与度至关重要。现有方法虽优化了知识推理或对话行为规划,但其与用户参与度的关系微妙,无法保证实际效果。为此,我们让交互式大模型通过对话未来发展的信号学习用户参与度,采用用户互动后对对话意图的反应作为奖励信号进行对齐。为此,我们构建用户模拟器,结合i×MCTS(交互式蒙特卡洛树搜索)探索用户与模型间的交互,生成高质量与低质量对话体验对。基于这些数据,使用直接偏好优化(DPO)对齐模型以实现高水平用户参与度。在情感支持对话和劝导性对话两个场景的实验表明,该方法能有效提升交互式大模型的用户参与度。

原文摘要 · Abstract (English)

Enhancing user engagement through interactions plays an essential role in socially-driven dialogues. While prior works have optimized models to reason over relevant knowledge or plan a dialogue act flow, the relationship between user engagement and knowledge or dialogue acts is subtle and does not guarantee user engagement in socially-driven dialogues. To this end, we enable interactive LLMs to learn user engagement by leveraging signals from the future development of conversations. Specifically, we adopt a more direct and relevant indicator of user engagement, i.e., the user's reaction related to dialogue intention after the interaction, as a reward to align interactive LLMs. To achieve this, we develop a user simulator to interact with target interactive LLMs and explore interactions between the user and the interactive LLM system via \textit{i$\times$MCTS} (\textit{M}onte \textit{C}arlo \textit{T}ree \textit{S}earch for \textit{i}nteraction). In this way, we collect a dataset containing pairs of higher and lower-quality experiences using \textit{i$\times$MCTS}, and align interactive LLMs for high-level user engagement by direct preference optimization (DPO) accordingly. Experiments conducted on two socially-driven dialogue scenarios (emotional support conversations and persuasion for good) demonstrate that our method effectively enhances user engagement in interactive LLMs.

对话系统用户参与大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。