用不精确的机器人创造有性格的艺术装置,让缺陷成为创意来源。
Semantic Glitch: Agency and Artistry in an Autonomous Pixel Cloud
- 用多模态大模型实现无传感器自主导航,靠语义理解而非精确感知。
- 13分钟飞行记录显示机器人行为不可预测但合理,体现个性特征。
- 适合对艺术机器人、人机交互、生成式智能感兴趣的读者。
主流机器人追求高精度与完美表现,本文探索了刻意‘低保真’方法的创作潜力。我们提出‘语义故障’(Semantic Glitch),一种以3D像素云为物理形态的自主飞行艺术装置,其形态源自数字考古学的‘物理故障’。我们设计了一种新颖的自主流程,摒弃传统传感器如LiDAR和SLAM,仅依赖多模态大语言模型进行定性语义理解来导航。通过自然语言提示为机器人构建类生物人格,形成‘叙事心智’,与‘弱化’且历史感厚重的实体相呼应。分析基于一段13分钟的自主飞行日志,并通过后续统计研究验证了该框架在量化生成不同人格方面的鲁棒性。综合分析揭示出涌现行为,包括基于地标导航、显著的‘计划-执行’差距,以及因缺乏精确本体感知而产生的不可预测却合理的性格特征。这证明了一种以不完美为特色的框架,其成功标准在于角色魅力而非效率。
原文摘要 · Abstract (English)
While mainstream robotics pursues metric precision and flawless performance, this paper explores the creative potential of a deliberately "lo-fi" approach. We present the "Semantic Glitch," a soft flying robotic art installation whose physical form, a 3D pixel style cloud, is a "physical glitch" derived from digital archaeology. We detail a novel autonomous pipeline that rejects conventional sensors like LiDAR and SLAM, relying solely on the qualitative, semantic understanding of a Multimodal Large Language Model to navigate. By authoring a bio-inspired personality for the robot through a natural language prompt, we create a "narrative mind" that complements the "weak," historically, loaded body. Our analysis begins with a 13-minute autonomous flight log, and a follow-up study statistically validates the framework's robustness for authoring quantifiably distinct personas. The combined analysis reveals emergent behaviors, from landmark-based navigation to a compelling "plan to execution" gap, and a character whose unpredictable, plausible behavior stems from a lack of precise proprioception. This demonstrates a lo-fi framework for creating imperfect companions whose success is measured in character over efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。