用合成数据训练轨迹描述模型,真实数据上表现优异
Text2Traj2Text: Learning-by-Synthesis Framework for Contextual Captioning of Human Movement Trajectories
- 用大模型生成真实感轨迹与情境描述数据
- 在真实人类轨迹上达到领先指标(ROUGE/BERT Score)
- 适合零售场景的顾客行为理解与智能推荐
本文提出Text2Traj2Text,一种基于合成学习的框架,用于为零售店中顾客的运动轨迹数据生成可能的情境描述。核心思想是利用大型语言模型在店铺地图上合成多样且逼真的情境描述及对应运动轨迹。尽管模型仅在完全合成的数据上训练,仍能良好泛化到真实人类参与者产生的轨迹与描述。系统性评估表明,该框架在ROUGE和BERT Score等指标上优于现有方法。
原文摘要 · Abstract (English)
This paper presents Text2Traj2Text, a novel learning-by-synthesis framework for captioning possible contexts behind shopper's trajectory data in retail stores. Our work will impact various retail applications that need better customer understanding, such as targeted advertising and inventory management. The key idea is leveraging large language models to synthesize a diverse and realistic collection of contextual captions as well as the corresponding movement trajectories on a store map. Despite learned from fully synthesized data, the captioning model can generalize well to trajectories/captions created by real human subjects. Our systematic evaluation confirmed the effectiveness of the proposed framework over competitive approaches in terms of ROUGE and BERT Score metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。