arXiv:2603.09992cs.CLcs.AI2026-03被引 1

为学术机构定制大模型对话系统,兼顾效果与责任部署。

TAMUSA-Chat: A Domain-Adapted Large Language Model Conversational System for Research and Responsible Deployment

  • 用监督微调+检索增强生成实现领域适配
  • 实证分析不同模型规模的训练效率与成本
  • 适合教育科研机构研究可信AI系统部署

本文提出TAMUSA-Chat,一个面向研究的领域自适应大语言模型对话系统框架。针对通用基础模型在机构场景中的适配难题,采用监督微调、检索增强生成及系统化评估方法。完整涵盖机构数据获取、预处理、嵌入构建、模型训练与部署策略,组件模块化支持配置复现。实证分析了不同模型规模与训练轮次下的微调行为,揭示领域适配效率、算力需求与质量-成本权衡。开源代码库(https://github.com/alsmadi/TAMUSA_LLM_Based_Chat_app)促进机构级LLM部署、评估方法与教育AI伦理研究。

原文摘要 · Abstract (English)

This paper presents TAMUSA-Chat, a research-oriented framework for building domain-adapted large language model conversational systems. The work addresses critical challenges in adapting general-purpose foundation models to institutional contexts through supervised fine-tuning, retrieval-augmented generation, and systematic evaluation methodologies. We describe the complete architecture encompassing data acquisition from institutional sources, preprocessing pipelines, embedding construction, model training workflows, and deployment strategies. The system integrates modular components enabling reproducible experimentation with training configurations, hyper-parameters, and evaluation protocols. Our implementation demonstrates how academic institutions can develop contextually grounded conversational agents while maintaining transparency, governance compliance, and responsible AI practices. Through empirical analysis of fine-tuning behavior across model sizes and training iterations, we provide insights into domain adaptation efficiency, computational resource requirements, and quality-cost trade-offs. The publicly available codebase at https://github.com/alsmadi/TAMUSA_LLM_Based_Chat_app supports continued research into institutional LLM deployment, evaluation methodologies, and ethical considerations for educational AI systems.

大模型对话系统机构部署负责任AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。