arXiv:2510.27126cs.HCcs.AI2025-10被引 2

用强化学习让问卷聊天机器人实时自适应,提升回答质量。

AURA: A Reinforcement Learning Framework for AI-Driven Adaptive Conversational Surveys

  • 基于强化学习动态选择提问方式,实时调整对话策略。
  • 响应质量提升0.076,指定问题减少63%,验证行为增加10倍。
  • 适合需要深度交互的调研场景,如校园氛围评估。

传统在线问卷个性化不足,导致参与度低、回答肤浅。尽管AI问卷聊天机器人提升了便利性,但多数仍为被动响应:依赖固定对话树或静态提示模板,无法在会话中动态适应个体用户,造成泛化追问和低质量回应。为此,我们提出AURA(Adaptive Understanding through Reinforcement Learning for Assessment),一种基于强化学习的自适应对话问卷框架。AURA采用四维LSDE指标(长度、自我披露、情绪、具体性)量化回答质量,并通过epsilon-greedy策略选择后续问题类型,实时更新每轮会话中的预期质量增益。系统初始化使用96个先前的校园氛围对话(共467次人机交互)提取先验知识,在10-15轮对话中平衡探索与利用,动态适应个体参与者。控制实验显示,AURA在响应质量上平均提升0.076,显著优于非自适应基线(p=0.044, d=0.66),主要得益于指定类问题减少63%及验证行为提升10倍。结果表明,强化学习可显著增强问卷机器人的自适应能力,将静态问卷转化为可交互、自优化的评估系统。

原文摘要 · Abstract (English)

Conventional online surveys provide limited personalization, often resulting in low engagement and superficial responses. Although AI survey chatbots improve convenience, most are still reactive: they rely on fixed dialogue trees or static prompt templates and therefore cannot adapt within a session to fit individual users, which leads to generic follow-ups and weak response quality. We address these limitations with AURA (Adaptive Understanding through Reinforcement Learning for Assessment), a reinforcement learning framework for AI-driven adaptive conversational surveys. AURA quantifies response quality using a four-dimensional LSDE metric (Length, Self-disclosure, Emotion, and Specificity) and selects follow-up question types via an epsilon-greedy policy that updates the expected quality gain within each session. Initialized with priors extracted from 96 prior campus-climate conversations (467 total chatbot-user exchanges), the system balances exploration and exploitation across 10-15 dialogue exchanges, dynamically adapting to individual participants in real time. In controlled evaluations, AURA achieved a +0.076 mean gain in response quality and a statistically significant improvement over non-adaptive baselines (p=0.044, d=0.66), driven by a 63% reduction in specification prompts and a 10x increase in validation behavior. These results demonstrate that reinforcement learning can give survey chatbots improved adaptivity, transforming static questionnaires into interactive, self-improving assessment systems.

对话系统强化学习问卷设计自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。