语音写作新工具,让移动中的人也能轻松写长文。
StepWrite: Adaptive Planning for Speech-Driven Text Generation
- 将写作拆解为步骤,用语音提示引导用户逐项完成。
- 实测显示认知负担降低,可用性和满意度显著提升。
- 适合开车、走路等双手忙碌场景下的长文本创作。
人们常通过语音转文字系统快速撰写短文本,但现有语音界面难以支持更复杂、需持续上下文追踪的长篇内容创作,尤其在移动过程中无法视觉查看进度时更为困难。长文本如结构化邮件或深度回复,需要持续跟踪上下文、结构化指导和对用户意图的动态适应,而传统语音输入与助手无法满足这些需求。本文提出 StepWrite,一个基于大语言模型的语音交互系统,通过分步引导、非视觉语音提示,实现移动中无需双手双眼的长文本创作。系统将写作过程分解为可管理子任务,动态根据上下文与用户意图调整提示,减轻用户记忆与规划负担。25名参与者在移动或手部占用状态下进行测试,结果表明其显著降低认知负荷,提升可用性与满意度。技术评估也验证了其在动态上下文提示生成、语气准确匹配和事实核查方面的有效性。本研究展示了结构化、上下文感知语音交互在日常多任务场景中提升免手持、免注视沟通的潜力。
原文摘要 · Abstract (English)
People frequently use speech-to-text systems to compose short texts with voice. However, current voice-based interfaces struggle to support composing more detailed, contextually complex texts, especially in scenarios where users are on the move and cannot visually track progress. Longer-form communication, such as composing structured emails or thoughtful responses, requires persistent context tracking, structured guidance, and adaptability to evolving user intentions--capabilities that conventional dictation tools and voice assistants do not support. We introduce StepWrite, a large language model-driven voice-based interaction system that augments human writing ability by enabling structured, hands-free and eyes-free composition of longer-form texts while on the move. StepWrite decomposes the writing process into manageable subtasks and sequentially guides users with contextually-aware non-visual audio prompts. StepWrite reduces cognitive load by offloading the context-tracking and adaptive planning tasks to the models. Unlike baseline methods like standard dictation features (e.g., Microsoft Word) and conversational voice assistants (e.g., ChatGPT Advanced Voice Mode), StepWrite dynamically adapts its prompts based on the evolving context and user intent, and provides coherent guidance without compromising user autonomy. An empirical evaluation with 25 participants engaging in mobile or stationary hands-occupied activities demonstrated that StepWrite significantly reduces cognitive load, improves usability and user satisfaction compared to baseline methods. Technical evaluations further confirmed StepWrite's capability in dynamic contextual prompt generation, accurate tone alignment, and effective fact checking. This work highlights the potential of structured, context-aware voice interactions in enhancing hands-free and eye-free communication in everyday multitasking scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。