arXiv:2604.15190cs.AIcs.CL2026-04中稿 · SIGIR 2026 Industr…

用双路径模拟提升商家策略评估准确率,避免过度理性化。

Meituan Merchant Business Diagnosis via Policy-Guided Dual-Process User Simulation

论文配图:Meituan Merchant Business Diagnosis via Policy-Guided Dual-Process User Simulation
图 1 · 摘自论文原文
  • 分两路模拟:大模型推理+机器学习拟合,互补纠错。
  • 在26000条轨迹上误差仅8.80%,优于基线45.8%以上。
  • 适合需要低成本评估商家策略的平台方使用。

模拟群体用户行为可实现无需昂贵线上实验的商家策略反事实评估。但构建可信模拟器面临两大结构性挑战:其一,信息不全导致基于推理的模拟器在缺乏线下情境和隐性习惯等未观测因素时过度理性化;其二,机制二元性要求同时捕捉可解释偏好与隐式统计规律,单一范式难以兼顾。为此提出政策引导的混合模拟(PGHS)框架,从行为轨迹中挖掘可迁移决策策略作为共享对齐层,锚定基于大模型的推理分支以防止过度理性化,同时接入基于机器学习的拟合分支以吸收隐式规律。两个分支的群体级预测结果融合,实现互补修正。在美团部署于101家商户及超过26,000条轨迹的数据上,PGHS实现8.80%的群体模拟误差,较最优推理与拟合基线分别降低45.8%和40.9%。

原文摘要 · Abstract (English)

Simulating group-level user behavior enables scalable counterfactual evaluation of merchant strategies without costly online experiments. However, building a trustworthy simulator faces two structural challenges. First, information incompleteness causes reasoning-based simulators to over-rationalize when unobserved factors such as offline context and implicit habits are missing. Second, mechanism duality requires capturing both interpretable preferences and implicit statistical regularities, which no single paradigm achieves alone. We propose Policy-Guided Hybrid Simulation (PGHS), a dual-process framework that mines transferable decision policies from behavioral trajectories and uses them as a shared alignment layer. This layer anchors an LLM-based reasoning branch that prevents over-rationalization and an ML-based fitting branch that absorbs implicit regularities. Group-level predictions from both branches are fused for complementary correction. We deploy PGHS on Meituan with 101 merchants and over 26,000 trajectories. PGHS achieves a group simulation error of 8.80%, improving over the best reasoning-based and fitting-based baselines by 45.8% and 40.9% respectively.

用户模拟策略评估双路径美团

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。