用可本地部署的因果大模型预测出行方式选择,效果超商用模型。
Towards Locally Deployable Fine-Tuned Causal Large Language Models for Mode Choice Behaviour
- 基于参数高效微调与损失掩码策略,定制专用因果大模型LiTransMC。
- F1得分0.6845,分布校准误差仅0.000245,优于大模型和传统方法。
- 适合交通政策研究、行为模拟及需要可解释性的本地化系统使用者。
本研究探索开放获取、可本地部署的因果大语言模型(LLM)在出行方式选择预测中的应用,提出首个为此任务微调的因果大模型LiTransMC。在三个陈述与揭示偏好数据集上,系统评测了11个开源大模型(参数规模1-12B),测试396种配置,生成超7.9万次出行选择决策。除预测精度外,采用BERTopic主题建模与新型解释强度指数评估模型推理过程,首次实现行为理论对齐的结构化分析。LiTransMC通过参数高效微调与损失掩码策略,取得加权F1分数0.6845与詹森-香农散度0.000245,超越未微调本地模型及更大规模专有系统(如GPT-4o),也优于离散选择模型与机器学习分类器。该模型兼具高瞬时准确率与近乎完美的分布校准能力,证明了专业化、可本地部署的因果大模型在预测与可解释性融合上的可行性。结合结构化行为预测与自然语言推理,为支持多任务、对话式交通建模、政策测试与行为洞察生成提供了新路径。研究为通用大模型向交通领域专用、可解释工具转化提供了可行方案,兼顾隐私保护、成本降低与访问普及。
原文摘要 · Abstract (English)
This study investigates the adoption of open-access, locally deployable causal large language models (LLMs) for travel mode choice prediction and introduces LiTransMC, the first fine-tuned causal LLM developed for this task. We systematically benchmark eleven open-access LLMs (1-12B parameters) across three stated and revealed preference datasets, testing 396 configurations and generating over 79,000 mode choice decisions. Beyond predictive accuracy, we evaluate models generated reasoning using BERTopic for topic modelling and a novel Explanation Strength Index, providing the first structured analysis of how LLMs articulate decision factors in alignment with behavioural theory. LiTransMC, fine-tuned using parameter efficient and loss masking strategy, achieved a weighted F1 score of 0.6845 and a Jensen-Shannon Divergence of 0.000245, surpassing both untuned local models and larger proprietary systems, including GPT-4o with advanced persona inference and embedding-based loading, while also outperforming classical mode choice methods such as discrete choice models and machine learning classifiers for the same dataset. This dual improvement, i.e., high instant-level accuracy and near-perfect distributional calibration, demonstrates the feasibility of creating specialist, locally deployable LLMs that integrate prediction and interpretability. Through combining structured behavioural prediction with natural language reasoning, this work unlocks the potential for conversational, multi-task transport models capable of supporting agent-based simulations, policy testing, and behavioural insight generation. These findings establish a pathway for transforming general purpose LLMs into specialized and explainable tools for transportation research and policy formulation, while maintaining privacy, reducing cost, and broadening access through local deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。