arXiv:2608.01366cs.MAcs.AI2026-08

用8个智能体对话引导用户优化提问,让大模型输出更精准。

Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution

论文配图:Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution
图 1 · 摘自论文原文
  • 八智能体系统通过问答逐步构建结构化提示
  • 用户提示完整度从42%提升至91%,单次完成任务
  • 适合需要高质量大模型输出的研究者与开发者

大型语言模型在复杂任务中至关重要,但输出质量受限于用户输入的提示。迭代式多轮提示常导致上下文退化和认知回报递减。我们提出PAWNI(基于神经智能的提示架构助手),一个由八个智能体组成的对话接口,通过自演化知识库指导的问答对话,将非结构化查询转化为结构化提示。不同于优化模型输出,PAWNI聚焦于优化问题本身,提前明确意图。我们还设计了包含18个提示要素的三层框架:基础、增强与提升。为评估系统行为并验证测量协议,我们开展了一项探索性被试内研究(N=4),覆盖四项复杂任务,整合32通道脑电图、NASA-TLX工作量量表及行为指标。使用PAWNI后,用户提示结构完整度达42%至91%,对大模型输出质量评分更高,工作负荷显著降低(NASA-TLX 39.6 vs. 21.7)。所有参与者均在单轮内获得满意结果,而自主操作需1-12轮。尽管效应大小因样本量小而不稳定,但方向一致性支持‘前置优化提示构造是人机协作关键杠杆’的假设。

原文摘要 · Abstract (English)

Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts. Iterative multi-turn prompting often leads to context degradation and diminishing cognitive returns. We present PAWNI (Prompt Architecture Wizard using Neural Intelligence), an agentic conversational interface of eight agents that transforms unstructured queries into structured prompts through guided question-and-answer dialogue informed by a self-evolving knowledge base. Rather than optimising the model's response, PAWNI optimises the question itself by front-loading intent clarification. We also propose a three-tier framework of 18 prompt elements across Essential, Enhancement, and Elevation categories. To evaluate system behaviour and validate a measurement protocol, we conducted an exploratory within-subjects study (N=4) across four complex tasks, integrating 32-channel EEG, NASA-TLX workload, and behavioural metrics. Participants produced more structurally complete prompts with PAWNI (42% to 91% of assessed elements), rated LLM outputs higher across all quality dimensions, and reported lower workload (39.6 vs. 21.7 NASA-TLX). Every participant reached satisfactory output in a single turn, compared to 1-12 turns unaided. While effect sizes are unstable due to sample size, direction consistency supports the hypothesis that optimising prompt formulation front-end is a critical lever for human-AI collaboration.

人机协作提示工程智能体系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。