用多智能体流程整合出行数据收集、建模与天气敏感需求预测。
An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction

- 设计三智能体系统:对话收集、数据处理、行为预测协同工作。
- 视觉增强模型达71.5%准确率,优于纯文本零样本模型的69.9%。
- 习惯性出行信息提升预测效果,适合交通规划与智能系统研究者。
出行行为研究日益融合数字数据采集与预测建模,但两者常独立开发与评估。本研究提出三智能体工作流,整合对话式数据采集、结构化数据处理与行为预测。通过聊天机器人实施、图像增强的显性偏好调查,在五个预设天气场景下收集学生通勤者的出行方式选择,共获得454个受试者-场景观测值。采用多元逻辑回归分析天气关联性,以逻辑回归和随机森林作为机器学习基准。评估了九个本地部署的大语言模型(参数量2至350亿),覆盖四种零样本提示与上下文条件,并拓展至角色设定、少样本及基于视觉的配置。随机森林在五分类任务中达到69.6%准确率,最佳纯文本零样本模型达69.9%且无需任务微调。习惯性出行信息带来最稳定提升,专家框架整体优于角色扮演,当缺乏习惯信息时角色设定最有用。少样本提示提升多个模型表现,增益在少量示例后趋于稳定。使用与受访者相同的天气图像,最优视觉配置达71.5%五分类准确率,表明视觉上下文可为部分模型提供额外预测信息。研究表明,对话调查、结构化处理、传统建模、机器学习与多模态大模型可在可审计的多智能体流程中有效协同。
原文摘要 · Abstract (English)
Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction. A chatbot-administered, image-augmented stated-preference survey collected mode choices from student commuters across five predefined weather scenarios, yielding 454 respondent-scenario observations. Weather-related associations were analyzed using a multinomial logit model, while logistic regression and random forest provided machine-learning benchmarks. Nine locally deployed large language models (LLMs), ranging from 2 to 35 billion parameters, were evaluated across four zero-shot prompt-and-context conditions and extended through persona, few-shot, and vision-based configurations. Random forest achieved 69.6% five-class accuracy, while the best text-only zero-shot LLM reached 69.9% without task-specific fitting. Habitual travel information produced the most consistent gains, Expert framing generally outperformed Role-Play, and persona information was most useful when habitual travel information was unavailable. Few-shot prompting improved prediction for several models, with gains stabilizing after a small number of examples. Using the same weather images shown to respondents, the best vision-based configuration reached 71.5% five-class accuracy, indicating that visual context may provide additional predictive information for selected models. Overall, the study shows how conversational surveys, structured data processing, conventional behavioral modeling, machine learning, and multimodal LLM prediction can be coordinated within an auditable multi-agent workflow.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。