arXiv:2603.24936cs.CVcs.AI2026-03

用行为规则引导轨迹生成,让智能系统更懂人类行走规律。

TIGFlow-GRPO: Trajectory Forecasting via Interaction-Aware Flow Matching and Reward-Guided Optimization

  • 分两阶段生成:先建交互图模型捕捉人与环境关系,再用奖励优化探索合理路径。
  • 在ETH/UCY和SDD数据集上,长时预测更稳定,社交合规性提升12.3%。
  • 适合自动驾驶、人群监控等需理解人类行为的复杂视觉场景应用。

人类轨迹预测对自动驾驶、人群监控等视觉复杂环境中的智能多媒体系统至关重要。尽管条件流匹配(CFM)在从时空观测中建模轨迹分布方面表现强劲,但现有方法仍以监督拟合为主,难以充分反映社会规范与场景约束。为此,我们提出TIGFlow-GRPO,一种两阶段生成框架,将基于流的轨迹生成与行为规则对齐。第一阶段构建基于CFM的预测器,引入轨迹-交互图(TIG)模块,精细建模人与人、人与场景的交互,增强上下文编码能力,为后续对齐提供更丰富的条件特征。第二阶段进行流-强化学习优化(Flow-GRPO)后训练,将确定性流展开重构成随机的ODE-to-SDE采样,实现轨迹探索;复合奖励同时考虑视域感知的社会合规性与地图感知的物理可行性。通过评估SDE采样探索出的轨迹,GRPO逐步引导多模态预测向行为合理未来收敛。在ETH/UCY和SDD数据集上的实验表明,TIGFlow-GRPO提升了预测精度与长时稳定性,生成轨迹更具社会合规性与物理可行性。结果表明,该方法有效实现了基于流的轨迹建模与行为感知对齐的结合。

原文摘要 · Abstract (English)

Human trajectory forecasting is important for intelligent multimedia systems operating in visually complex environments, such as autonomous driving and crowd surveillance. Although Conditional Flow Matching (CFM) has shown strong ability in modeling trajectory distributions from spatio-temporal observations, existing approaches still focus primarily on supervised fitting, which may leave social norms and scene constraints insufficiently reflected in generated trajectories. To address this issue, we propose TIGFlow-GRPO, a two-stage generative approach that aligns flow-based trajectory generation with behavioral rules. In the first stage, we build a CFM-based predictor with a Trajectory-Interaction-Graph (TIG) module to model fine-grained visual-spatial interactions and strengthen context encoding. This stage captures both agent-agent and agent-scene relations more effectively, providing more informative conditional features for subsequent alignment. In the second stage, we perform Flow-GRPO post-training, where deterministic flow rollout is reformulated as stochastic ODE-to-SDE sampling to enable trajectory exploration, and a composite reward combines view-aware social compliance with map-aware physical feasibility. By evaluating trajectories explored through SDE rollout, GRPO progressively steers multimodal predictions toward behaviorally plausible futures. Experiments on the ETH/UCY and SDD datasets show that TIGFlow-GRPOimproves forecasting accuracy and long-horizon stability while generatingtrajectories that are more socially compliant and physically feasible.These results suggest that the proposed approach provides an effective way to connectflow-based trajectory modeling with behavior-aware alignment in dynamic multimedia environments.

轨迹预测流模型行为对齐强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。