Responsive image
博碩士論文 etd-0526125-194853 詳細資訊
Title page for etd-0526125-194853
論文名稱
Title
投影導向之軟提示微調:邁向可解釋且語義對齊之表示學習
Projection-Guided Soft Prompt Tuning: Toward Interpretable and Semantically Aligned Representations
系所名稱
Department
畢業學年期
Year, semester
語文別
Language
學位類別
Degree
頁數
Number of pages
63
研究生
Author
指導教授
Advisor
召集委員
Convenor
口試委員
Advisory Committee
口試日期
Date of Exam
2025-06-03
繳交日期
Date of Submission
2025-06-26
關鍵字
Keywords
軟提示微調、可解釋性、高維嵌入更新、離散化、投影
Soft Prompt Tuning, Interpretability, High-dimensional Embedding Update, Discretization, Projection
統計
Statistics
本論文已被瀏覽 219 次,被下載 0
The thesis/dissertation has been browsed 219 times, has been downloaded 0 times.
中文摘要
軟提示微調(Soft Prompt Tuning)方法作為一種高效的參數微調技術,能在保留預訓練模型參數的情況下,透過調整少量提示向量,保有預訓練模型能力的同時,高效達到特定任務的模型表現。然而,現有方法雖能提升模型表現,卻缺乏對軟提示嵌入語義變化的有效解釋,目前大多研究往往僅聚焦於模型輸出的預測結果,忽略了嵌入向量間的語義演變過程,使得訓練過程中的語義轉換無法被具體呈現,為解決此問題,本研究提出一種具備離散化嵌入空間投影的軟提示微調方法,以兼顧解釋性與模型表現。
大多研究中的離散化方法通常直接將軟提示嵌入向量替換為語義鄰近的離散化嵌入,雖可增強解釋性,但容易導致模型性能大幅下降,具體原因為直接替換忽略了嵌入向量間的細微語義變化,且此離散化得到的嵌入,其語義往往與任務無關。為改善此問題,本研究提出一種嵌入空間投影策略,以投影參與更新方式進行語義聚合,在提升解釋性的同時減少對模型性能的負面影響。
研究結果顯示,經過離散化嵌入空間投影後的軟提示在多個文本分類任務上展現出較佳的語義對齊性,並且透過鄰近詞語生成與視覺化分析,逐漸揭示訓練過程中的語義集中趨勢,同時,大幅減緩了模型表現的損害。綜合來看,本研究在語義解釋性與模型表現之間取得了有效平衡,驗證了我們方法的可行性與有效性。
Abstract
Soft Prompt Tuning, as an efficient parameter adjustment technique, can effectively enhance model performance on specific tasks while retaining the pretrained model's capabilities by adjusting a small number of prompt vectors. However, existing methods, although capable of improving model performance, lack a comprehensive explanation of the semantic evolution of prompt embeddings. Most studies primarily focus on the prediction results of the model output, overlooking the semantic transformation among embedding vectors during training. This omission hinders the concrete interpretation of the training process. To address this issue, this study proposes a Soft Prompt Tuning method based on discrete embedding space projection, aiming to balance interpretability and model performance.
In conventional methods, the discrete approach typically replaces the prompt embeddings directly with semantically similar discrete embeddings, which may enhance interpretability but often leads to significant performance degradation. This is primarily due to the direct replacement disregarding subtle semantic variations among embeddings, and the resulting discrete embeddings are often unrelated to the task semantics. To mitigate this issue, this study introduces an embedding space projection strategy, wherein prompt embeddings participate in the projection update process to achieve semantic aggregation, thereby enhancing interpretability while minimizing the negative impact on model performance.
The experimental results demonstrate that the proposed method exhibits superior semantic alignment in multiple text classification tasks after the discrete embedding space projection process. Additionally, through the generation of neighboring words and visualization analysis, the semantic concentration trend during training is progressively revealed, significantly alleviating the performance degradation. Overall, the proposed method effectively balances semantic interpretability and model performance, validating its feasibility and effectiveness.
目次 Table of Contents
論文審定書 .............................................................................................................................. i
摘要 ......................................................................................................................................... ii
Abstract ................................................................................................................................... iii
目錄 ......................................................................................................................................... v
圖次 ....................................................................................................................................... vii
表次 ...................................................................................................................................... viii
第一章 緒論 ............................................................................................................................. 1
1.1 研究背景 ......................................................................................................................... 1
1.2 研究動機 ......................................................................................................................... 2
1.3 研究目的 ......................................................................................................................... 2
第二章 文獻探討 ..................................................................................................................... 4
2.1 提示微調 ......................................................................................................................... 4
2.1.1 硬提示和軟提示 ...................................................................................................... 4
2.2 軟提示微調 ..................................................................................................................... 6
2.2.1 軟提示微調相關方法 .............................................................................................. 7
2.3 軟提示的解釋性 ............................................................................................................. 8
第三章 研究方法 .................................................................................................................... 1
3.1 研究方法架構 ............................................................................................................... 13
3.2 標準軟提示微調階段 ................................................................................................... 16
3.3 投影優化階段 ............................................................................................................... 18
3.3.1 訓練資料構建的嵌入空間 .................................................................................... 19
3.3.2 投影方法與優化策略 ............................................................................................ 21
3.4 解釋性設計 ................................................................................................................... 23
第四章 實驗 ........................................................................................................................... 26
4.1 資料集 ........................................................................................................................... 26
4.2 基線(Baselines) ............................................................................................................. 29
4.3 實驗設計 ....................................................................................................................... 31
4.4 實驗結果 ....................................................................................................................... 32
4.4.1 模型表現評估 ........................................................................................................ 34
4.4.2 可解釋性評估 ........................................................................................................ 39
4.4.2.1 軌跡視覺化分析 ............................................................................................. 40
4.4.2.2 語義鄰近詞分析 ............................................................................................. 42
4.4.2.3 Token Entropy ................................................................................................. 43
4.5 綜合討論 ....................................................................................................................... 49
第五章 結論 ........................................................................................................................... 51
5.1 結論 ............................................................................................................................... 51
5.2 未來展望 ....................................................................................................................... 51
參考文獻 ................................................................................................................................ 53
參考文獻 References
參考文獻

[1] B. Lester, R. Al-Rfou, and N. Constant, "The power of scale for parameter-efficient prompt tuning," arXiv preprint arXiv:2104.08691, 2021.
[2] T. Schick and H. Schütze, "It's not just size that matters: Small language models are also few-shot learners," arXiv preprint arXiv:2009.07118, 2020.
[3] P. Passigan, K. Yohannes, and J. Pereira, "Continuous Prompt Generation from Linear Combination of Discrete Prompt Embeddings," arXiv preprint arXiv:2312.10323, 2023.
[4] Q. Guo et al., "Connecting large language models with evolutionary algorithms yields powerful prompt optimizers," arXiv preprint arXiv:2309.08532, 2023.
[5] R. Pryzant, D. Iter, J. Li, Y. T. Lee, C. Zhu, and M. Zeng, "Automatic prompt optimization with" gradient descent" and beam search," arXiv preprint arXiv:2305.03495, 2023.
[6] D. Khashabi et al., "Prompt waywardness: The curious case of discretized interpretation of continuous prompts," arXiv preprint arXiv:2112.08348, 2021.
[7] T. Ju, Y. Zheng, H. Wang, H. Zhao, and G. Liu, "Is continuous prompt a combination of discrete prompts? towards a novel view for interpreting continuous prompts," in Findings of the Association for Computational Linguistics: ACL 2023, 2023, pp. 7804-7819.
[8] O. Patel, J. Wang, N. S. Nayak, S. Srinivas, and H. Lakkaraju, "Towards Interpretable Soft Prompts," arXiv preprint arXiv:2504.02144, 2025.
[9] Y. Wen, N. Jain, J. Kirchenbauer, M. Goldblum, J. Geiping, and T. Goldstein, "Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery," Advances in Neural Information Processing Systems, vol. 36, 2024.
[10] F. Doshi-Velez and B. Kim, "Towards a rigorous science of interpretable machine learning," arXiv preprint arXiv:1702.08608, 2017.
[11] T. B. Brown, "Language models are few-shot learners," arXiv preprint arXiv:2005.14165, 2020.
[12] Q. Chen et al., "Lifelong knowledge editing for llms with retrieval-augmented continuous prompt learning," arXiv preprint arXiv:2405.03279, 2024.
[13] X. L. Li and P. Liang, "Prefix-tuning: Optimizing continuous prompts for generation," arXiv preprint arXiv:2101.00190, 2021.
[14] A. Razdaibiedina, Y. Mao, R. Hou, M. Khabsa, M. Lewis, and A. Almahairi, "Progressive prompts: Continual learning for language models," arXiv preprint arXiv:2301.12314, 2023.
[15] D. Wingate, M. Shoeybi, and T. Sorensen, "Prompt compression and contrastive conditioning for controllability and toxicity reduction in language models," arXiv preprint arXiv:2210.03162, 2022.
[16] L. K. Şenel, I. Utlu, V. Yücesoy, A. Koc, and T. Cukur, "Semantic structure and interpretability of word embeddings," IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 26, no. 10, pp. 1769-1779, 2018.
[17] M. Wistuba, P. T. Sivaprasad, L. Balles, and G. Zappella, "Choice of peft technique in continual learning: Prompt tuning is not all you need," arXiv preprint arXiv:2406.03216, 2024.
[18] T. Zhang, C. Zhang, J. X. Morris, E. Bagdasarian, and V. Shmatikov, "Soft prompts go hard: Steering visual language models with hidden meta-instructions," arXiv preprint arXiv:2407.08970, 2024.
[19] R. Belanec, S. Ostermann, I. Srba, and M. Bielikova, "Task prompt vectors: Effective initialization through multi-task soft-prompt transfer," arXiv preprint arXiv:2408.01119, 2024.
[20] A. K. Mohankumar, P. Nema, S. Narasimhan, M. M. Khapra, B. V. Srinivasan, and B. Ravindran, "Towards transparent and explainable attention models," arXiv preprint arXiv:2004.14243, 2020.

電子全文 Fulltext
本電子全文僅授權使用者為學術研究之目的,進行個人非營利性質之檢索、閱讀、列印。請遵守中華民國著作權法之相關規定,切勿任意重製、散佈、改作、轉貼、播送,以免觸法。
論文使用權限 Thesis access permission:自定論文開放時間 user define
開放時間 Available:
校內 Campus:開放下載的時間 available 2028-06-26
校外 Off-campus:開放下載的時間 available 2028-06-26

您的 IP(校外) 位址是 18.97.9.175
現在時間是 2026-08-11
論文校外開放下載的時間是 2028-06-26

Your IP address is 18.97.9.175
The current date is 2026-08-11
This thesis will be available to you on 2028-06-26.

紙本論文 Printed copies
紙本論文的公開資訊在102學年度以後相對較為完整。如果需要查詢101學年度以前的紙本論文公開資訊,請聯繫圖資處紙本論文服務櫃台。如有不便之處敬請見諒。
開放時間 available 2028-06-26

QR Code