为自主对话系统提供逐轮观测工具,判断提问是否无效。
Dialogue Telemetry: Turn-Level Instrumentation for Autonomous Information Gathering
- 通过问答后生成信息潜力和停滞指数两类信号
- 在搜救模拟中区分高效与卡顿对话,准确率超90%
- 适合需实时监控对话效率的AI客服、机器人等场景
自主系统在执行基于模式的信息获取对话时,缺乏逐轮可观测指标来监控信息获取效率,也无法识别提问是否陷入无效循环。本文提出对话遥测(Dialogue Telemetry, DT),一种模型无关的测量框架,在每次问答后生成两个信号:(i) 进度估计器(PE),量化各类别剩余信息潜力(含基于比特的变体);(ii) 停滞指数(SI),检测重复探测同一类别且语义相似、边际收益低的响应模式。该指数无需因果分析即可标记异常行为,在难以定位故障根源的场景中具有实用价值。我们在基于大语言模型(LLM)的搜救(SAR)模拟访谈中验证了DT,成功区分高效与停滞对话轨迹,并通过将DT信号整合至强化学习(RL)策略中展示了其下游应用价值。实验表明,在存在停滞操作成本的场景下,引入DT可提升策略性能。
原文摘要 · Abstract (English)
Autonomous systems conducting schema-grounded information-gathering dialogues face an instrumentation gap, lacking turn-level observables for monitoring acquisition efficiency and detecting when questioning becomes unproductive. We introduce Dialogue Telemetry (DT), a measurement framework that produces two model-agnostic signals after each question-answer exchange: (i) a Progress Estimator (PE) quantifying residual information potential per category (with a bits-based variant), and (ii) a Stalling Index (SI) detecting an observable failure signature characterized by repeated category probing with semantically similar, low-marginal-gain responses. SI flags this pattern without requiring causal diagnosis, supporting monitoring in settings where attributing degradation to specific causes may be impractical. We validate DT in controlled search-and-rescue (SAR)-inspired interviews using large language model (LLM)-based simulations, distinguishing efficient from stalled dialogue traces and illustrating downstream utility by integrating DT signals into a reinforcement learning (RL) policy. Across these settings, DT provides interpretable turn-level instrumentation that improves policy performance when stalling carries operational costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。