用图文混合提示让大模型理解软件工程图示
The importance of visual modelling languages in generative software engineering
- 结合图表与自然语言作为输入,驱动大模型完成软件工程任务
- 首次探索图文混合提示在软件工程中的实际应用案例
- 适合关注AI辅助编程与可视化建模的研究者和开发者
多模态GPT代表了软件工程与生成式人工智能之间互动的转折点。GPT-4不仅接受自然语言输入,还能处理图像和文本。我们研究了由此带来的相关应用场景。据我们所知,尚无其他工作探究过利用图文混合提示通过多模态GPT完成软件工程任务的使用案例。
原文摘要 · Abstract (English)
Multimodal GPTs represent a watershed in the interplay between Software Engineering and Generative Artificial Intelligence. GPT-4 accepts image and text inputs, rather than simply natural language. We investigate relevant use cases stemming from these enhanced capabilities of GPT-4. To the best of our knowledge, no other work has investigated similar use cases involving Software Engineering tasks carried out via multimodal GPTs prompted with a mix of diagrams and natural language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。