arXiv:2508.03583cs.MMcs.IR2025-08被引 5

构建1.4万条多模态生活日志问答数据集,支持记忆辅助等个性化应用。

OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset

  • 基于18个月可穿戴设备采集的多模态日志数据构建问答对。
  • 包含14,187组问答,覆盖多种问题类型与难度等级。
  • 适合研究个人助理、健康支持和生活方式分析的开发者。

我们提出了OpenLifelogQA,一个由18个月多模态生活日志数据构建的大规模开放式问答数据集。生活日志指通过可穿戴设备被动收集并分析个人日常活动,生成包括图像、位置和生物特征在内的丰富多模态数据。对生活日志进行问答可使用户交互式查询自身经历,支持记忆辅助、生活方式分析和个人协助等应用。OpenLifelogQA包含14,187个问答对,涵盖多种问题类型与难度层次,旨在支持真实场景下的稳健评估。相比已有资源,该数据集在多样性和实用性上更具优势。为建立基线,我们评估了LLaVA-NeXT-Interleave 7B模型,取得89.7% BERTScore、25.87% ROUGE-L和平均LLM Score 3.97的成绩。通过发布OpenLifelogQA,我们旨在推动生活日志技术的未来研究,助力实现具备记忆增强、健康支持和生活方式指导能力的个人助手。

原文摘要 · Abstract (English)

We introduce OpenLifelogQA, a large-scale open-ended lifelog QA dataset constructed from 18 months of multimodal lifelog data. Lifelogging is the passive collection and analysis of personal daily activities using wearable devices, producing rich multimodal data such as images, locations, and biometrics. Question answering (QA) over lifelog data enables users to interactively query their own experiences, supporting applications in memory support, lifestyle analysis, and personal assistance. OpenLifelogQA contains 14,187 Q&A pairs spanning multiple question types and difficulty levels, designed to support robust evaluation in realistic settings. Compared with prior resources, OpenLifelogQA offers greater diversity and practicality for real-world applications. To establish baselines, we evaluate the LLaVA-NeXT-Interleave 7B model, achieving 89.7% BERTScore, 25.87% ROUGE-L, and an average LLM Score of 3.97. By releasing OpenLifelogQA, we aim to promote future research on lifelog technologies, paving the way for personal lifelog assistants capable of memory augmentation, healthcare support, and lifestyle coaching.

生活日志多模态问答系统个人助理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。