通过逆向推理构建用户视角的结构化表示,解决隐私数据稀缺问题。
Situation Graph Prediction for User Perspective Modeling
- 将用户视角建模转化为从多模态痕迹反推内在状态的逆问题。
- 三款前沿模型均显示隐状态推理显著难于表层信息提取。
- 采用结构优先合成数据策略,无需真实标注即可训练与验证。
情境感知的人工智能需要建模动态演变的内在状态——目标、情绪、上下文,而不仅仅是偏好。进展受限于数据瓶颈:数字足迹具有隐私敏感性,且视角状态极少被标注。我们提出情境图预测(SGP)任务,将用户视角建模视为一个逆向推理问题:从可观测的多模态痕迹重构结构化、符合本体论的视角表示,适合作为个人代理的长期记忆。为在无真实标签情况下实现可落地,我们采用结构优先的合成生成策略,通过设计使潜在标签与可观测痕迹对齐。作为初步探索,我们构建了一个数据集,并使用检索增强的上下文学习作为监督代理进行诊断研究。在三个前沿基础模型(GPT-4o、Gemini 2.5 Flash、Claude Sonnet 4)上,我们观察到表面信息提取与潜在视角推断之间存在一致的正向差距,表明在受控环境下,潜在状态推理始终比表面提取更困难。结果表明SGP任务具有挑战性,同时支持结构优先数据合成策略的有效性。为保障可复现性与透明度,代码、数据及补充资源已公开:https://github.com/flybits/sg-prediction。
原文摘要 · Abstract (English)
Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data bottleneck: digital footprints are privacy-sensitive and perspective states are rarely labeled. We propose Situation Graph Prediction (SGP), a task that frames user perspective modeling as an inverse inference problem: reconstructing structured, ontology-aligned representations of perspective from observable multimodal artifacts, suitable as long-horizon memory for personal agents. To enable grounding without real labels, we use a structure-first synthetic generation strategy that aligns latent labels and observable traces by design. As a pilot, we construct a dataset and run a diagnostic study using retrieval-augmented in-context learning as a proxy for supervision. In our diagnostic study across three frontier foundation models (GPT-4o, Gemini 2.5 Flash, Claude Sonnet 4), we observe a consistent positive gap between surface-level extraction and latent perspective inference---indicating that latent-state inference is consistently harder than surface extraction under our controlled setting. Results suggest SGP is non-trivial and provide evidence for the structure-first data synthesis strategy. For reproducibility and transparency, our code, data, and supplementary resources are available here: https://github.com/flybits/sg-prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。