用音频和视觉信号让机器人识别家中突发状况并及时响应。
HomeEmergency -- Using Audio to Find and Respond to Emergencies in the Home
- 构建多模态场景图,通过贝叶斯推理定位声音来源
- 利用视觉语言模型判断物品危险属性,识别紧急事件
- 在消费级机器人上验证方法可迁移性,适合家庭安全研究
美国每年意外居家死亡超过12.8万例。本文旨在使家用机器人能够识别并响应家庭紧急情况,减少伤亡。我们基于ThreeDWorld仿真器构建了一个新的家庭紧急事件数据集,每个场景以瞬时或周期性声音为起点,可能代表紧急状况。机器人需结合先前观察、音频与图像信息,在多房间环境中判断是否存在紧急情况。我们提出一种模块化方法,核心是新颖的概率动态场景图(P-DSG),其中代理节点通过概率边表示,经贝叶斯推断后实现高效定位。同时引入多模态视觉语言模型(VLMs)分析物体属性(如可燃性)并识别紧急事件。我们在消费级机器人上展示了该方法在真实任务中的可行性,证明了任务与方法的可迁移性。数据集将在论文发表后公开。
原文摘要 · Abstract (English)
In the United States alone accidental home deaths exceed 128,000 per year. Our work aims to enable home robots who respond to emergency scenarios in the home, preventing injuries and deaths. We introduce a new dataset of household emergencies based in the ThreeDWorld simulator. Each scenario in our dataset begins with an instantaneous or periodic sound which may or may not be an emergency. The agent must navigate the multi-room home scene using prior observations, alongside audio signals and images from the simulator, to determine if there is an emergency or not. In addition to our new dataset, we present a modular approach for localizing and identifying potential home emergencies. Underpinning our approach is a novel probabilistic dynamic scene graph (P-DSG), where our key insight is that graph nodes corresponding to agents can be represented with a probabilistic edge. This edge, when refined using Bayesian inference, enables efficient and effective localization of agents in the scene. We also utilize multi-modal vision-language models (VLMs) as a component in our approach, determining object traits (e.g. flammability) and identifying emergencies. We present a demonstration of our method completing a real-world version of our task on a consumer robot, showing the transferability of both our task and our method. Our dataset will be released to the public upon this papers publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。