MUSE让用户像操控机器人一样直观理解并纠正大模型数据科学流程中的错误。
MUSE: An Interactive Meta-Agent for Understanding and Steering LLM-powered Data Science Systems

- 将执行日志按语义分层,支持从宏观到细节的自由切换浏览。
- 用户可直接针对步骤提问、反馈或修改,无需手动查找历史记录。
- 自动识别可疑步骤并辅助修复,实现人机协同决策。
大型语言模型的进步催生了新型智能数据科学系统,用户可通过自然语言完成复杂数据分析任务。尽管大幅减少人工操作,但当系统出错或输出异常时,仍难以诊断行为并干预推理过程。我们提出 MUSE,一种交互式元代理,通过(1)将底层执行轨迹动态重构为多层级语义结构,支持从高层概览到低层实现细节的导航;(2)允许用户在上下文中引用特定工作流步骤,提出有根据的问题、提供反馈并修正问题步骤,而无需手动定位执行历史;(3)通过提示可疑步骤、支撑修复流程,并将用户修复意图转化为上下文相关的指令,实现混合主动性引导。一项包含15名参与者的对照实验表明,MUSE显著提升了任务效率,并增强了用户对智能数据科学流程的理解与控制信心。
原文摘要 · Abstract (English)
Recent advances in large language models have enabled a new class of agentic data science systems that allow users to complete complex data science workflows through natural language. Although these systems can significantly reduce manual effort, it remains difficult to diagnose their behavior and steer the reasoning process when failures or unexpected outputs occur. We present MUSE, an interactive meta-agent that enhances user understanding and control of agentic data science systems by (1) dynamically restructuring low-level execution traces into multiple semantic levels that support navigation from high-level overviews to low-level implementation details; (2) enabling users to reference specific workflow steps in context to ask grounded questions, provide feedback, and revise problematic steps without manually locating relevant execution history; and (3) supporting mixed-initiative steering by surfacing suspicious steps for inspection, scaffolding the repair process, and translating user repair intent into contextualized instructions for the underlying agent. In a between-subjects study (n = 15), MUSE improved task efficiency and increased users' confidence in understanding and steering agentic data science workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。