评测大模型在带噪声工具下的对话决策能力,发现现有模型无法有效提示教师避免依赖错误建议。
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

- 构建教育场景下带噪声工具的对话评估基准,模拟真实决策环境
- 提出新指标RAIR/RSR,衡量多轮对话中人与AI的依赖程度变化
- 实测13个主流模型均在噪声环境下暴露信任误判问题,适合教育类AI系统研究者
大型语言模型正被广泛应用于医疗、教育、金融等高风险领域的任务导向型对话系统。然而现有基准通常假设工具输出完全准确,忽略了实际部署中工具存在噪声及人类决策者对代理信任度不确定的现实。我们引入PREDACTBENCH,一个基于教育场景的评估基准,利用可验证的真实结果与明确干预决策作为测试基础。首先,构建了支持人工决策的AI辅助对话基准,其中AI使用有噪声的预测器辅助用户。其次,提出回合级相对AI依赖度(RAIR)和相对自依赖度(RSR)指标,扩展了以往的信任校准框架至多轮对话。第三,在两个教育数据集上评估13个最先进的闭源与开源大模型:OULAD(英国开放大学的真实评估轨迹)和PREDACT-CS(60门课程,包含真实最终成绩与合成的每周分数轨迹),并开展教师与助教的人类实验。结果显示,当工具存在噪声时,当前顶级模型未能为教师提供足够可见性,导致其可能过度依赖错误建议或幻觉内容。我们公开PREDACTBENCH,以促进更可靠的AI决策支持系统在教育中的发展。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in task-oriented dialogue systems that support multi-step decision-making in high-stakes domains such as education, healthcare, and finance. However, existing benchmarks typically assume perfectly accurate tool outputs, overlooking the reality that deployed systems must operate with noisy tools and human decision-makers whose trust in the agent is itself uncertain. Such conditions are common in practice, for example, a clinician using a diagnostic prediction tool or an advisor relying on a model that forecasts student outcomes from historical records. We introduce PREDACTBENCH, a benchmark for evaluating dialogue agents paired with statistically imperfect tools, using education as a measurable testbed where ground truth outcomes and clear intervention decisions are available. First, we build a benchmark for AI-assisted human decision-making, where the AI uses noisy predictors to help guide a user. Second, we introduce episode-level Relative AI-Reliance (RAIR) and Relative self-reliance (RSR) metrics, extending prior trust calibration framework to multi-turn dialogue. Third, we evaluate 13 state-of-the-art closed and open source LLMs on two educational datasets, OULAD (real assessment trajectories from the UK Open University) and PREDACT-CS (60 courses with real final grade outcomes and synthetically generated weekly score trajectories), alongside a human study with instructors and teaching assistants. We find that when tools are noisy, SOTA models are supposed to provide visibility to teachers so that they do not over-rely on wrong suggestions or hallucinations, but current models fail to do that. We offer PREDACTBENCH to help build better LLMs as AI decision support systems to help teachers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。