智能判断何时提问,提升纠正健康谣言的效率与效果。
Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention

- 根据用户认知水平和信念程度,动态决定是否提问以获取更多信息
- 在三个数据集上比始终提问的基线少用30%对话轮次,效果更优
- 适合需要精准干预的医疗问答、人机对话系统开发者
纠正对话中的健康谣言不仅需要提供事实反驳,还需考虑用户的知识水平、信念状态和信息需求。有效的干预常需先提出合适的澄清问题。然而现有方法要么立即回应,要么盲目追问,未区分提问的价值与成本。本文提出奖励优化型探问-回应框架(RO-PnR),在每一轮对话中权衡提问带来的预期收益与交互成本,决定是继续探问还是直接给出修正。通过为模拟用户建模健康素养与信念坚定度的潜在状态,捕捉用户异质性对提问价值的影响。实验表明,在三个健康谣言数据集和三种基础模型上,RO-PnR在成本调整后的效用最高,平均比始终探问的基线减少30%的对话轮次。
原文摘要 · Abstract (English)
Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and what they need to hear, so an effective intervention often depends on first asking the right clarifying question. Yet existing methods either respond immediately or probe indiscriminately, treating clarification as either unnecessary or always beneficial. We propose Reward-Optimized Probe-and-Respond (RO-PnR), a framework that learns when asking is worth its cost. At each turn, RO-PnR chooses between probing for more information and committing to a final correction, guided by a turn-level reward that weighs the expected gain from probing against its interaction cost. To capture how user heterogeneity affects probing value, we model each simulated user with a latent state along health literacy and belief commitment. Experiments show that RO-PnR achieves the highest cost-adjusted utility across three health-misinformation datasets and three base models, using 30% fewer turns than always-probe baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。