用大模型增强贝叶斯网络,从少量出行调查数据生成更真实的合成数据。
LEBGen: An LLM-Enhanced Bayesian Network Framework for Few-Shot Travel Survey Data Generation

- 结合大模型识别出行人群画像,优化贝叶斯网络结构
- 在2%小样本下,分布和依赖关系误差显著降低
- 适合交通规划与行为分析中数据稀缺场景
出行调查数据对交通规划与行为分析至关重要,但大规模代表性样本采集成本高、耗时长。一种实用替代方案是从少量样本生成合成记录。然而,小样本难以覆盖多元出行群体,且不足以揭示人口特征与出行行为间的复杂依赖关系。现有方法各有局限:贝叶斯网络(BN)可显式控制分布,但小样本学习的结构可能遗漏真实依赖或保留虚假关联;大语言模型(LLMs)可补充行为知识,弥补统计证据不足。为此,本文提出LEBGen——一种基于大模型增强的贝叶斯网络框架,利用LLM从人口属性与出行行为统计中识别出行人群画像,恢复被忽略的依赖关系并剔除虚假关联,再仅基于观测数据参数化优化后的网络以生成合成记录。在2022年香港出行特征调查的2%小样本设置下,LEBGen将均值边缘Jensen-Shannon散度从0.0671降至0.0091,平均绝对Cramer's V误差比最优基线降低14.3%,显著提升分布与依赖保真度。
原文摘要 · Abstract (English)
Travel survey data are essential for transportation planning and travel behavior analysis, yet collecting large-scale representative samples is costly and time-consuming. A practical alternative is to generate synthetic survey records from a few-shot sample. However, such samples provide incomplete coverage of heterogeneous traveler groups and insufficient evidence for recovering the complex dependencies between demographic characteristics and travel behavior. Existing approaches have complementary limitations. Probabilistic generative models such as Bayesian networks (BNs) offer explicit distributional control, but structures learned from few-shot samples may omit meaningful dependencies or retain spurious ones. Large language models (LLMs) can help address these difficulties in BN structure learning by providing behavioral knowledge that complements the limited statistical evidence. We therefore propose LEBGen, an LLM-enhanced BN framework that uses this knowledge to refine network structure for few-shot travel survey data generation. Specifically, the LLM first identifies traveler personas from demographic attribute and travel behavior statistics, then recovers dependencies missed by the persona-augmented BN structure and prune spurious ones. The refined BN is parameterized exclusively from the observed data to generate synthetic records. Under a 2% few-shot setting on the 2022 Hong Kong Travel Characteristics Survey, LEBGen reduces the mean marginal Jensen-Shannon divergence from 0.0671 to 0.0091 and the mean absolute Cramer's V error by 14.3% over the best-performing baseline, substantially improving both distributional and dependency fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。