MuRAL数据集为智能家居中的多居民活动提供自然语言标注,助力大模型理解日常行为。
MuRAL: A Multi-Resident Ambient Sensor Dataset Annotated with Natural Language for Activities of Daily Living
- 构建21小时多居民传感器数据,含自然语言描述与身份标签
- 大模型在长期身份追踪、动作描述和上下文推理上仍存明显短板
- 适合研究智能环境中的多用户行为建模与大模型应用
大型语言模型(LLMs)在基于环境传感器数据的人类活动理解中展现出先进推理与零样本识别能力。然而,现有广泛使用的多居民数据集如CASAS、ARAS和MARBLE缺乏自然语言上下文和细粒度标注,限制了大模型在真实智能环境中的充分应用。为此,我们提出了MuRAL(多居民环境传感器数据集,附带自然语言标注),包含来自21个会话的超过21小时多用户传感器数据。该数据集独特地融合了详细的自然语言描述、明确的居民身份标识以及丰富的活动标签,均置于复杂动态的多居民场景中。我们在MuRAL上对前沿大模型进行了三项核心任务的基准测试:主体归属、动作描述与活动分类。结果显示,当前大模型在长序列中的准确居民分配、精确动作描述生成以及上下文有效整合方面仍面临重大挑战。数据集已公开:https://mural.imag.fr/
原文摘要 · Abstract (English)
Recent progress in Large Language Models (LLMs) has enabled advanced reasoning and zero-shot recognition for human activity understanding with ambient sensor data. However, widely used multi-resident datasets such as CASAS, ARAS, and MARBLE lack natural language context and fine-grained annotation, limiting the full exploitation of LLM capabilities in realistic smart environments. To address this gap, we present MuRAL (Multi-Resident Ambient sensor dataset with natural Language), comprising over 21 hours of multi-user sensor data from 21 sessions in a smart home. MuRAL uniquely features detailed natural language descriptions, explicit resident identities, and rich activity labels, all situated in complex, dynamic, multi-resident scenarios. We benchmark state-of-the-art LLMs on MuRAL for three core tasks: subject assignment, action description, and activity classification. Results show that current LLMs still face major challenges on MuRAL, especially in maintaining accurate resident assignment over long sequences, generating precise action descriptions, and effectively integrating context for activity prediction. The dataset is publicly available at: https://mural.imag.fr/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。