arXiv:2602.23366cs.HCcs.IR2026-02

用可交互的模块化转换器,让AI帮你跨文档、跨格式智能改写文档。

Doc To The Future: Infomorphs for Interactive, Multimodal Document Transformation and Generation

  • 提出'信息形态'(infomorph)概念,支持用户控制多模态内容转换流程。
  • 通过可视化编排工具,实现文档提取、摘要、重排版等操作的灵活组合。
  • 适合需要精细控制AI生成过程的研究者和知识工作者使用。

从现有文档中合成新文档是多个领域知识工作的重要环节,通常涉及从多份文档中收集内容、组织信息,并转化为报告、幻灯片或表格等形式。尽管生成式AI在自动化部分流程上展现出潜力,但其对多模态输入输出的控制能力有限。本文提出“infomorph”概念——一种模块化、用户可操控的AI增强型信息转换机制,支持可控的信息融合与跨格式重构。我们构建了一个设计空间,结合生成式AI与用户意图,实现灵活、交互式、多模态的文档创作。作为具体实现,我们推出DocuCraft,一个基于画布的界面,允许用户可视化编排infomorph工作流,执行页面提取、内容摘要、格式转换与生成等操作,全程利用生成式AI支持跨文档、跨模态的丰富变换。通过实例演示,展示了DocuCraft在常见知识工作场景中的流畅协作能力,凸显了生成式AI辅助信息工作的透明性与模块化交互潜力。

原文摘要 · Abstract (English)

Creating new documents by synthesizing information from existing sources is an important part of knowledge work in many domains. This process often involves gathering content from multiple documents, organizing it, and then transforming it into new forms such as reports, slides, or spreadsheets. While recent advances in Generative AI have shown potential in automating parts of this process, they often provide limited user control over the handling of multimodal inputs and outputs. In this work, we introduce the notion of "infomorphs" which are modular, user-steerable, AI-augmented transformations that support controlled synthesis, and restructuring of information across formats and modalities. We propose a design space that leverage infomorph-driven workflows to enable flexible, interactive, and multimodal document creation by combining Generative AI techniques with user intent and desired information context. As a concrete instantiation of this design space, we present DocuCraft, a canvas-based interface to visually compose infomorph workflows. DocuCraft allows users to chain together infomorphs that perform operations such as page extraction, content summarization, reformatting, and generation, leveraging Generative AI at each stage to support rich, cross-document and cross-modal transformations. We demonstrate the capabilities of DocuCraft through an example-driven usage scenario that spans across different facets of common knowledge work tasks illustrating its support for fluid, human-in-the-loop document synthesis and highlights opportunities for more transparent and modular interaction for Generative AI-assisted information work.

文档生成人机协作生成式AI多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。