大模型靠少量例子就能预测人一天的活动和时长,适合数据少的智能环境应用。
Evaluating Few-Shot Temporal Reasoning of LLMs for Human Activity Prediction in Smart Environments
- 用四种上下文信息增强提示,让大模型从少量例子中推理人类行为
- 零样本也能生成合理日程,加一两个例子显著提升时间估算准确度
- 适合智能家居、人机协作等低数据场景,可替代传统数据驱动模型
预测人类活动及其持续时间在智能家居自动化、基于仿真的建筑设计、基于活动的交通系统模拟及人机协作等领域至关重要,要求自适应系统能响应人类行为。现有数据驱动的基于代理模型(从规则系统到深度学习)在低数据环境下表现不佳,限制了实际应用。本文探讨预训练于广泛人类知识的大语言模型是否可通过紧凑上下文线索推理日常活动,填补这一空白。采用检索增强提示策略,融合时间、空间、行为历史与人格特征四类上下文,在CASAS Aruba智能家庭数据集上评估。测试涵盖两个互补任务:带持续时间估计的下一步活动预测,以及多步日程生成,每项任务在不同数量的少样本示例下进行评估。分析少样本效应揭示了在数据效率与预测准确性之间取得平衡所需的上下文监督量,尤其在低数据环境中。结果表明,大语言模型具备强大的内在时间理解能力:即使在零样本设置下,仍能生成连贯的每日活动预测;增加一两个示范后,持续时间校准和类别准确度进一步提升。超过少数示例后性能趋于饱和,显示边际收益递减。序列级评估确认了不同少样本条件下的一致时间对齐。研究结果表明,预训练语言模型可作为有前景的时间推理器,既能捕捉重复性习惯,也能反映情境依赖的行为变化,从而强化基于代理模型中的行为模块。
原文摘要 · Abstract (English)
Anticipating human activities and their durations is essential in applications such as smart-home automation, simulation-based architectural and urban design, activity-based transportation system simulation, and human-robot collaboration, where adaptive systems must respond to human activities. Existing data-driven agent-based models--from rule-based to deep learning--struggle in low-data environments, limiting their practicality. This paper investigates whether large language models, pre-trained on broad human knowledge, can fill this gap by reasoning about everyday activities from compact contextual cues. We adopt a retrieval-augmented prompting strategy that integrates four sources of context--temporal, spatial, behavioral history, and persona--and evaluate it on the CASAS Aruba smart-home dataset. The evaluation spans two complementary tasks: next-activity prediction with duration estimation, and multi-step daily sequence generation, each tested with various numbers of few-shot examples provided in the prompt. Analyzing few-shot effects reveals how much contextual supervision is sufficient to balance data efficiency and predictive accuracy, particularly in low-data environments. Results show that large language models exhibit strong inherent temporal understanding of human behavior: even in zero-shot settings, they produce coherent daily activity predictions, while adding one or two demonstrations further refines duration calibration and categorical accuracy. Beyond a few examples, performance saturates, indicating diminishing returns. Sequence-level evaluation confirms consistent temporal alignment across few-shot conditions. These findings suggest that pre-trained language models can serve as promising temporal reasoners, capturing both recurring routines and context-dependent behavioral variations, thereby strengthening the behavioral modules of agent-based models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。