arXiv:2511.05817cs.HCcs.MM2025-11中稿 · AAAI被引 5

用语音+手绘实时生成设计草图,让创意更流畅。

TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech

  • 结合手绘与语音输入,实现边画边说的创作方式
  • 用户测试显示文本提示会打断创意思维流
  • 适合需要快速发散创意的设计师或团队

草图是早期设计构思中广泛使用的表达方式。尽管生成式AI聊天机器人在创意生成中日益普及,但设计师常难以构建有效提示,且仅靠文字难以表达不断演变的视觉概念。在一项包含6名设计师的形成性研究中,我们发现基于文本的提示会破坏创作流程。为此,我们开发了TalkSketch——一个嵌入式的多模态AI草图系统,融合自由手绘与实时语音输入。该系统通过捕捉绘制过程中的口头描述,生成上下文感知的AI响应,旨在支持更流畅的设计构思过程。研究强调,生成式AI应融入设计过程本身,而非仅关注最终输出。

原文摘要 · Abstract (English)

Sketching is a widely used medium for generating and exploring early-stage design concepts. While generative AI (GenAI) chatbots are increasingly used for idea generation, designers often struggle to craft effective prompts and find it difficult to express evolving visual concepts through text alone. In the formative study (N=6), we examined how designers use GenAI during ideation, revealing that text-based prompting disrupts creative flow. To address these issues, we developed TalkSketch, an embedded multimodal AI sketching system that integrates freehand drawing with real-time speech input. TalkSketch aims to support a more fluid ideation process through capturing verbal descriptions during sketching and generating context-aware AI responses. Our work highlights the potential of GenAI tools to engage the design process itself rather than focusing on output.

多模态生成设计工具语音交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。