用一周第一人称视频训练ChatGPT,看它能否读懂我的生活
Can ChatGPT Learn My Life From a Week of First-Person Video?
- 用54小时第一视角视频生成多级摘要,微调GPT-4o系列模型
- 模型准确识别出居住地、身份、惯用手和养猫等关键信息
- 但会虚构人物姓名,存在明显幻觉问题
受生成式AI和可穿戴摄像设备(如智能眼镜、AI纽扣)进步的推动,本文探究基础模型通过第一人称摄像头数据学习佩戴者个人生活的能力。研究者连续一周佩戴摄像头,累计拍摄54小时,生成了分钟级、小时级和天级等多层级摘要,并对GPT-4o与GPT-4o-mini进行微调。通过查询微调后的模型,可反推其学到的内容。结果显示:两模型均掌握了基本个人信息(如大致年龄、性别);其中GPT-4o还正确推断出研究者居住于匹兹堡,是卡内基梅隆大学博士生,右利手,且养有一只猫。但两者均存在严重幻觉,会虚构视频中人物的名字。
原文摘要 · Abstract (English)
Motivated by recent improvements in generative AI and wearable camera devices (e.g. smart glasses and AI-enabled pins), I investigate the ability of foundation models to learn about the wearer's personal life through first-person camera data. To test this, I wore a camera headset for 54 hours over the course of a week, generated summaries of various lengths (e.g. minute-long, hour-long, and day-long summaries), and fine-tuned both GPT-4o and GPT-4o-mini on the resulting summary hierarchy. By querying the fine-tuned models, we are able to learn what the models learned about me. The results are mixed: Both models learned basic information about me (e.g. approximate age, gender). Moreover, GPT-4o correctly deduced that I live in Pittsburgh, am a PhD student at CMU, am right-handed, and have a pet cat. However, both models also suffered from hallucination and would make up names for the individuals present in the video footage of my life.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。