Responsive image
博碩士論文 etd-0526125-195046 詳細資訊
Title page for etd-0526125-195046
論文名稱
Title
基於專家混合架構的高效微調與輕量化模型的終身學習研究
Efficient Fine-Tuning and Lightweight Models for Lifelong Learning Based on MoE Architecture
系所名稱
Department
畢業學年期
Year, semester
語文別
Language
學位類別
Degree
頁數
Number of pages
56
研究生
Author
指導教授
Advisor
召集委員
Convenor
口試委員
Advisory Committee
口試日期
Date of Exam
2025-06-03
繳交日期
Date of Submission
2025-06-26
關鍵字
Keywords
終身學習、持續學習、專家混合模型、自然語言處理、深度學習
Lifelong Learning, Continual Learning, Mixture of Experts, Natural Language Processing, Deep Learning
統計
Statistics
本論文已被瀏覽 344 次,被下載 0
The thesis/dissertation has been browsed 344 times, has been downloaded 0 times.
中文摘要
在自然語言處理任務持續增長與語言模型參數量不斷增長的情況下,如何實現模型的終身學習能力並兼顧效率,是目前重要的研究方向之一。傳統的多任務學習雖能進行知識共享,卻依賴完整資料做重新訓練,難以應對真實場景中的任務序列與知識保留。本研究提出一套基於混合專家架構(Mixture of Experts, MoE)的終身學習策略,設計兩種正交方法:OrthoRouter與OrthoExpert,分別針對專家選擇策略與參數更新機制進行優化,以緩解終身學習常見的災難性遺忘問題,並提升模型資源使用效率。
本研究以T5-large為基礎,將前饋層替換為MoE結構,在OrthoRouter方法上,結合LoRA實現低參數微調,並進一步導入正交路由機制,使路由器根據輸入特徵選擇與既有知識正交的專家進行處理。而在OrthoExpert方法上,於專家層新增 LoRA 參數時實施正交投影與後續微調時增加懲罰項,確保參數學習不干擾舊任務知識。實驗設計採用經典NLP任務資料集,評估模型在三組任務順序下的平均準確率與記憶保持能力。
結果顯示,OrthoExpert在所有任務序列中均顯著優於現有的先進方法(如O-LoRA、LFPT5),並展現穩定的抗遺忘能力與推理效率。OrthoRouter同樣提升模型在多任務序列情況下的性能表現,且在專家負載均衡與專家分化方面效果明顯。此外,本研究架構在微調與推理階段皆能降低顯存消耗與參數啟用量,提升整體運算效率。
綜合而言,本研究所提出的「基於專家混合架構的高效微調與輕量化模型的終身學習研究」,不僅有效緩解知識遺忘問題,亦具大幅降低所需參數量與顯存消耗,為未來多任務與資源受限環境中的語言模型應用提供一種新的可能性。
Abstract
With the continuous expansion of natural language processing (NLP) tasks and the growing scale of language models, how to enable models to perform lifelong learning while maintaining both efficiency and accuracy has become a critical research challenge. Although traditional multi-task learning allows for knowledge sharing, it relies heavily on full data retraining, making it inadequate for real-world task sequences and memory retention. This study proposes a lifelong learning framework based on the Mixture of Experts (MoE) architecture and introduces two orthogonality-based strategies—OrthoRouter and OrthoExpert—to optimize expert selection and parameter update mechanisms. These strategies aim to mitigate catastrophic forgetting and improve resource efficiency.
Built upon the T5-large model, this research replaces the feedforward layers with a MoE structure. In the OrthoRouter approach, LoRA is used for parameter-efficient fine-tuning, and an orthogonal routing mechanism is introduced to dynamically select experts that are orthogonal to existing knowledge. In the OrthoExpert approach, newly added LoRA parameters in each expert are projected onto a space orthogonal to existing parameters, and an orthogonality penalty is applied during fine-tuning to prevent interference with previously learned knowledge. Experiments are conducted on benchmark NLP datasets, evaluating the average accuracy and memory retention across three different task sequences.
Experimental results show that OrthoExpert significantly outperforms state-of-the-art methods such as O-LoRA and LFPT5 across all task orders, demonstrating strong resistance to forgetting and efficient inference. OrthoRouter also enhances model performance in multi-task sequences, particularly in expert load balancing and specialization. Furthermore, the proposed architecture effectively reduces GPU memory usage and active parameter count during both fine-tuning and inference stages, contributing to improved computational efficiency.
In summary, this study presents a lifelong learning framework for efficient fine-tuning and lightweight modeling based on expert mixture architectures, which not only effectively mitigates catastrophic forgetting but also achieves high parameter efficiency and scalability. It offers a promising solution for deploying language models in multi-task and resource-constrained environments.
目次 Table of Contents
論文審定書 i
摘要 ii
Abstract iii
圖次 viii
表次 ix
第一章 緒論 1
1.1 研究背景 1
1.2 研究動機 2
1.2.1 終身學習於自然語言處理領域 2
1.2.2 混合專家帶來的高效推理應用於終身學習 3
1.3 研究目的 4
第二章 文獻探討 6
2.1 終身學習(Continual Learning, CL) 6
2.1.1 基於正則化的方法(Regularization-based Approach) 6
2.1.2 基於重放的方法(Replay-based Approach) 7
2.1.3 基於優化的方法(Optimization-based Approach) 8
2.1.4 基於表示的方法(Representation-based Approach) 8
2.1.5 基於架構的方法(Architecture-based Approach) 9
2.2 混合專家架構(MoE, Mixture of Experts) 9
2.2.1 稀疏MoE 9
2.2.2 MoE在Transformer 和大型語言模型中的應用 10
2.2.3 MoE應用於終身學習 10
第三章 研究方法與步驟 12
3.1 模型設計 12
3.2 正交MoE方法 13
3.2.1 LoRA Expert 14
3.2.2 Router正交選擇機制(OrthoRouter) 15
3.2.3 專家參數正交化(OrthoExpert) 18
第四章 實驗設置 22
4.1 資料集介紹 22
4.2 比較基準 23
4.3 實作細節 25
4.3.1 任務排序 25
4.3.2 超參數設定 25
4.3.3 專家數量 26
4.3.4 硬體設備 26
4.4 評估方式 27
第五章 實驗結果與分析 28
5.1 實驗結果 28
5.2 災難性遺忘分析 30
5.2.1 平均準確率分析 30
5.2.2 單任務記憶分析 31
5.3 消融實驗 33
5.3.1 MoE架構 33
5.3.2 OrthoRouter 34
5.3.3 OrthoExpert 35
5.4 效率分析 36
5.5 專家平衡分析 37
第六章 結論與未來展望 39
6.1 結論 39
6.2 未來展望 40
參考文獻 41
參考文獻 References
參考文獻
[1] H. Touvron et al., "Llama: Open and efficient foundation language models," arXiv preprint arXiv:2302.13971, 2023.
[2] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "Bert: Pre-training of deep bidirectional transformers for language understanding," in Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), 2019, pp. 4171-4186.
[3] C. Raffel et al., "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer," (in English), Journal of Machine Learning Research, vol. 21, no. 140, pp. 1-67, 2020. [Online]. Available: <Go to ISI>://WOS:000558791600001.
[4] OpenAI. "Openai: Introducing chatgpt." https://openai.com/ (accessed.
[5] B. Min et al., "Recent advances in natural language processing via large pre-trained language models: A survey," ACM Computing Surveys, vol. 56, no. 2, pp. 1-40, 2023.
[6] T. Wu, L. Luo, Y.-F. Li, S. Pan, T.-T. Vu, and G. Haffari, "Continual learning for large language models: A survey," arXiv preprint arXiv:2402.01364, 2024.
[7] J. Zheng, S. Qiu, C. Shi, and Q. Ma, "Towards lifelong learning of large language models: A survey," ACM Computing Surveys, vol. 57, no. 8, pp. 1-35, 2025.
[8] L. Wang, X. Zhang, H. Su, and J. Zhu, "A Comprehensive Survey of Continual Learning: Theory, Method and Application," IEEE Trans Pattern Anal Mach Intell, vol. 46, no. 8, pp. 5362-5383, Aug 2024, doi: 10.1109/TPAMI.2024.3367329.
[9] J. Kirkpatrick et al., "Overcoming catastrophic forgetting in neural networks," Proc Natl Acad Sci U S A, vol. 114, no. 13, pp. 3521-3526, Mar 28 2017, doi: 10.1073/pnas.1611835114.
[10] F. Zenke, B. Poole, and S. Ganguli, "Continual learning through synaptic intelligence," in International conference on machine learning, 2017: PMLR, pp. 3987-3995.
[11] A. Chaudhry, P. K. Dokania, T. Ajanthan, and P. H. Torr, "Riemannian walk for incremental learning: Understanding forgetting and intransigence," in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 532-547.
[12] X. Liu, M. Masana, L. Herranz, J. Van de Weijer, A. M. Lopez, and A. D. Bagdanov, "Rotate your networks: Better weight consolidation and less catastrophic forgetting," in 2018 24th International Conference on Pattern Recognition (ICPR), 2018: IEEE, pp. 2262-2268.
[13] L. Caccia, E. Belilovsky, M. Caccia, and J. Pineau, "Online learned continual compression with adaptive quantization modules," in International conference on machine learning, 2020: PMLR, pp. 1240-1250.
[14] A. Van Den Oord and O. Vinyals, "Neural discrete representation learning," Advances in neural information processing systems, vol. 30, 2017.
[15] S. Hou, X. Pan, C. C. Loy, Z. Wang, and D. Lin, "Learning a unified classifier incrementally via rebalancing," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 831-839.
[16] F. Mi, L. Chen, M. Zhao, M. Huang, and B. Faltings, "Continual Learning for Natural Language Generation in Task-oriented Dialog Systems," in Findings of the Association for Computational Linguistics: EMNLP 2020, T. Cohn, Y. He, and Y. Liu, Eds., November 2020, Online: Association for Computational Linguistics, pp. 3461–3474.
[17] H. Shin, J. K. Lee, J. Kim, and J. Kim, "Continual learning with deep generative replay," Advances in neural information processing systems, vol. 30, 2017.
[18] F.-K. Sun, C.-H. Ho, and H.-Y. Lee, "LAMOL: LAnguage MOdeling for Lifelong Language Learning," in International Conference on Learning Representations, 2020.
[19] X. Liu et al., "Generative feature replay for class-incremental learning," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 226-227.
[20] A. Iscen, J. Zhang, S. Lazebnik, and C. Schmid, "Memory-efficient incremental learning through feature adaptation," in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVI 16, 2020: Springer, pp. 699-715.
[21] K. Zhu, W. Zhai, Y. Cao, J. Luo, and Z.-J. Zha, "Self-sustaining representation expansion for non-exemplar class-incremental learning," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 9296-9305.
[22] E. Belouadah and A. Popescu, "Il2m: Class incremental learning with dual memory," in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 583-592.
[23] Z. Gong, K. Zhou, W. X. Zhao, J. Sha, S. Wang, and J.-R. Wen, "Continual pre-training of language models for math problem understanding with syntax-aware memory network," in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 5923-5933.
[24] D. Lopez-Paz and M. A. Ranzato, "Gradient episodic memory for continual learning," Advances in neural information processing systems, vol. 30, 2017.
[25] A. Chaudhry, M. A. Ranzato, M. Rohrbach, and M. Elhoseiny, "Efficient Lifelong Learning with A-GEM," in International Conference on Learning Representations, 2019.
[26] M. Farajtabar, N. Azizan, A. Mott, and A. Li, "Orthogonal gradient descent for continual learning," in International Conference on Artificial Intelligence and Statistics, 2020: PMLR, pp. 3762-3773.
[27] X. Wang et al., "Orthogonal Subspace Learning for Language Model Continual Learning," in Findings of the Association for Computational Linguistics: EMNLP 2023, 2023, pp. 10658-10671.
[28] R. Han, X. Ren, and N. Peng, "ECONET: Effective Continual Pretraining of Language Models for Event Temporal Reasoning," in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 5367-5380.
[29] Q. Liu, O. Majumder, A. Achille, A. Ravichandran, R. Bhotika, and S. Soatto, "Incremental few-shot meta-learning via indirect discriminant alignment," in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VII 16, 2020: Springer, pp. 685-701.
[30] Z. Wang et al., "Meta-learning with less forgetting on large-scale non-stationary task distributions," in European Conference on Computer Vision, 2022: Springer, pp. 221-238.
[31] A. Mallya, D. Davis, and S. Lazebnik, "Piggyback: Adapting a single network to multiple tasks by learning to mask weights," in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 67-82.
[32] A. Mallya and S. Lazebnik, "Packnet: Adding multiple tasks to a single network by iterative pruning," in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018, pp. 7765-7773.
[33] R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, "Adaptive Mixtures of Local Experts," Neural Comput, vol. 3, no. 1, pp. 79-87, Spring 1991, doi: 10.1162/neco.1991.3.1.79.
[34] M. I. Jordan and R. A. Jacobs, "Hierarchical Mixtures of Experts and the Em Algorithm," (in English), Neural Computation, vol. 6, no. 2, pp. 181-214, Mar 1994, doi: DOI 10.1162/neco.1994.6.2.181.
[35] N. Shazeer et al., "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer," in International Conference on Learning Representations, 2017.
[36] W. Fedus, B. Zoph, and N. Shazeer, "Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity," (in English), Journal of Machine Learning Research, vol. 23, no. 120, pp. 1-39, 2022. [Online]. Available: <Go to ISI>://WOS:001003360300001.
[37] D. Lepikhin et al., "GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding," in International Conference on Learning Representations, 2021.
[38] N. Du et al., "Glam: Efficient scaling of language models with mixture-of-experts," in International Conference on Machine Learning, 2022: PMLR, pp. 5547-5569.
[39] M. J. Jung and J. Kim, "PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning," arXiv preprint arXiv:2407.21571, 2024.
[40] S. Yang, M. A. Ali, C.-L. Wang, L. Hu, and D. Wang, "MoRAL: MoE Augmented LoRA for LLMs' Lifelong Learning," arXiv preprint arXiv:2402.11260, 2024.
[41] W. Chen et al., "Lifelong language pretraining with distribution-specialized experts," in International Conference on Machine Learning, 2023: PMLR, pp. 5383-5395.
[42] H. Li, S. Lin, L. Duan, Y. Liang, and N. B. Shroff, "Theory on mixture-of-experts in continual learning," arXiv preprint arXiv:2406.16437, 2024.
[43] Z. Li and D. Hoiem, "Learning without forgetting," IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 12, pp. 2935-2947, 2017.
[44] Z. Wang et al., "Learning to prompt for continual learning," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 139-149.
[45] C. Qin and S. Joty, "LFPT5: A Unified Framework for Lifelong Few-shot Language Learning Based on Prompt Tuning of T5," in International Conference on Learning Representations, 2022.
電子全文 Fulltext
本電子全文僅授權使用者為學術研究之目的,進行個人非營利性質之檢索、閱讀、列印。請遵守中華民國著作權法之相關規定,切勿任意重製、散佈、改作、轉貼、播送,以免觸法。
論文使用權限 Thesis access permission:自定論文開放時間 user define
開放時間 Available:
校內 Campus:開放下載的時間 available 2028-06-26
校外 Off-campus:開放下載的時間 available 2028-06-26

您的 IP(校外) 位址是 18.97.9.175
現在時間是 2026-08-11
論文校外開放下載的時間是 2028-06-26

Your IP address is 18.97.9.175
The current date is 2026-08-11
This thesis will be available to you on 2028-06-26.

紙本論文 Printed copies
紙本論文的公開資訊在102學年度以後相對較為完整。如果需要查詢101學年度以前的紙本論文公開資訊,請聯繫圖資處紙本論文服務櫃台。如有不便之處敬請見諒。
開放時間 available 2028-06-26

QR Code