論文使用權限 Thesis access permission:自定論文開放時間 user define
開放時間 Available:
校內 Campus:開放下載的時間 available 2027-06-17
校外 Off-campus:開放下載的時間 available 2027-06-17
論文名稱 Title |
利用知識圖擴充進行生物醫學問答系統效能與效率之研究 Research on Performance and Efficiency of Biomedical Question-Answering Systems through Knowledge Graph Expansion |
||
系所名稱 Department |
|||
畢業學年期 Year, semester |
語文別 Language |
||
學位類別 Degree |
頁數 Number of pages |
66 |
|
研究生 Author |
|||
指導教授 Advisor |
|||
召集委員 Convenor |
|||
口試委員 Advisory Committee |
|||
口試日期 Date of Exam |
2024-06-05 |
繳交日期 Date of Submission |
2024-06-17 |
關鍵字 Keywords |
知識圖、K-BERT、K-RET、四元架構、語言模型、醫學問答系統 Knowledge Graph, K-BERT, K-RET, quadruplet, language model, biomedical question-answering system |
||
統計 Statistics |
本論文已被瀏覽 417 次,被下載 0 次 The thesis/dissertation has been browsed 417 times, has been downloaded 0 times. |
中文摘要 |
在醫學領域,建立一個強大且高效的問答系統至關重要,它可以幫助醫護人員和患者獲取關於健康、疾病和治療的即時資訊。然而,醫學知識的快速增長和不斷變化,使得維護和更新傳統醫學問答系統變得極為困難。 一部分的學者開始研究如何透過重新預訓練大型語言模型來提升效能,但資源取得相當不易;另一派學者開始研究是否能在現有的大型語言模型基礎上,增加外來知識源以提供大型語言模型學習額外的知識來修正的答案讓答案可以與時俱進。 本研究希望利用目前既有的知識圖,並將之從三元結構擴充為四元的結構同時利用上述的K-BERT[1]之延伸K-RET[2]架構將四元結構的知識圖與現有語言模型如BioBERT[3]、SciBERT[4]等進行結合,在有限的運算資源下針對生物醫學問答系統進行效能的提升。 在實驗結果的部分,本研究得出的結論是使用結構資訊當作第四元進行三元組的篩選,確實能夠提升語言模型的效能,entity的使用量會有很明顯的減少,訓練時間是有機會降低的,同時本研究也驗證了使用第四元當作知識篩選的依據是可以有效降低訓練過程中所產生的雜訊的。 |
Abstract |
In the field of biomedical, establishing a robust and efficient question-answering system is crucial as it can assist healthcare professionals and patients in accessing real-time information about health, diseases, and treatments. However, the rapid growth and continuous evolution of medical knowledge make maintaining and updating traditional biomedical question-answering systems extremely challenging. Some scholars have begun to explore ways to enhance performance by retraining large-scale language models, but acquiring resources for this endeavor is quite difficult. Another group of researchers is investigating whether it is possible to augment existing large-scale language models with external knowledge sources to provide additional knowledge for refining answers and keeping them up-to-date. This study aims to utilize existing knowledge graphs and expand them from a ternary structure to a quaternary structure. Simultaneously, we aim to integrate this quaternary knowledge graph with existing language models such as BioBERT[3], SciBERT[4], etc., using the K-RET[2] framework, an extension of K-BERT[1]. This integration is intended to enhance the performance of medical question-answering systems within limited computational resources. In terms of experimental results, our conclusion is that using structure information as the fourth element for ternary tuple filtering indeed improves the performance of language models. There is a significant reduction in the usage of entities, potentially decreasing training time. Additionally, we have validated that using the fourth element as a basis for knowledge filtering effectively reduces the noise generated during the training process. |
目次 Table of Contents |
論文審定書 i 摘要 ii Abstract iii 目錄 iv 圖次 vii 表次 ix 第一章 緒論 1 1.1 研究背景 1 1.2 研究動機 1 1.3 研究目的 2 第二章 文獻探討 4 2.1 增強語言模型在特定領域之表現 4 2.1.1 ERNIE 5 2.1.2 MedBERT 7 2.1.3 DRAGON 8 2.1.4 K-BERT 10 2.1.5 K-RET 13 2.2 知識圖結構探討 15 2.2.1 四元結構之知識圖-Properties 15 2.2.2 四元結構之知識圖-Temporal 16 第三章 研究方法 18 3.1 資料前處理 19 3.1.1 醫學問答系統格式 19 3.1.2 訓練資料處理 19 3.1.3 知識圖處理 20 3.1.4 知識圖四元架構擴增 20 3.2 模型架構 23 3.2.1 句子樹(Sentence Tree) 24 3.2.2 Mask-Transformer 25 3.3 四元架構知識圖泛化性之研究(GENERALIZATION STUDY) 25 第四章 研究結果 26 4.1 實驗流程及設計 26 4.2 資料集及知識圖介紹 27 4.2.1 Medmcqa資料集 27 4.2.2 MedQA資料集 27 4.2.3 ChEBI知識圖 27 4.2.4 GO知識圖 28 4.3 評估指標 29 4.4 實驗參數設置 29 4.5 MEDMCQA實驗結果 30 4.5.1 三元知識圖測試 30 4.5.2 第四元測試 32 4.5.3 泛化性測試 34 4.5.4 GO知識圖四元驗證 37 4.5.5 小結 39 4.6 MEDQA實驗結果 39 4.6.1 三元知識圖測試 39 4.6.2 第四元測試 40 4.6.3 泛化性測試 43 4.6.4 GO知識圖四元驗證 45 4.6.5 小結 47 4.7 實驗結果探討 47 第五章 結論 51 5.1 結論 51 5.2 未來展望 52 參考文獻 53 |
參考文獻 References |
[1] W. Liu et al., "K-bert: Enabling language representation with knowledge graph," in Proceedings of the AAAI Conference on Artificial Intelligence, 2020, vol. 34, no. 03, pp. 2901-2908. [2] D. F. Sousa and F. M. Couto, "K-RET: knowledgeable biomedical relation extraction system," Bioinformatics, vol. 39, no. 4, 2023, doi: 10.1093/bioinformatics/btad174. [3] J. Lee et al., "BioBERT: a pre-trained biomedical language representation model for biomedical text mining," Bioinformatics, vol. 36, no. 4, pp. 1234-1240, 2020. [4] I. Beltagy, K. Lo, and A. Cohan, "SciBERT: A pretrained language model for scientific text," arXiv preprint arXiv:1903.10676, 2019. [5] "ChatGPT." https://chat.openai.com/ (accessed. [6] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," arXiv pre-print server, 2019-05-24 2019, doi: arxiv:1810.04805. [7] L. Li et al., "Real-world data medical knowledge graph: construction and applications," Artificial Intelligence in Medicine, vol. 103, p. 101817, 2020/03/01/ 2020, doi: https://doi.org/10.1016/j.artmed.2020.101817. [8] Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, "Xlnet: Generalized autoregressive pretraining for language understanding," Advances in neural information processing systems, vol. 32, 2019. [9] A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and Samuel, "GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding," arXiv pre-print server, 2019-02-22 2019, doi: arxiv:1804.07461. [10] Y. Peng, S. Yan, and Z. Lu, "Transfer learning in biomedical natural language processing: an evaluation of BERT and ELMo on ten benchmarking datasets," arXiv preprint arXiv:1906.05474, 2019. [11] Y. Gu et al., "Domain-specific language model pretraining for biomedical natural language processing," ACM Transactions on Computing for Healthcare (HEALTH), vol. 3, no. 1, pp. 1-23, 2021. [12] Z. Zhang, X. Han, Z. Liu, X. Jiang, M. Sun, and Q. Liu, "ERNIE: Enhanced language representation with informative entities," arXiv preprint arXiv:1905.07129, 2019. [13] A. Bordes, N. Usunier, A. Garcia-Duran, J. Weston, and O. Yakhnenko, "Translating embeddings for modeling multi-relational data," Advances in neural information processing systems, vol. 26, 2013. [14] P. Ferragina and U. Scaiella, "Tagme: on-the-fly annotation of short text fragments (by wikipedia entities)," in Proceedings of the 19th ACM international conference on Information and knowledge management, 2010, pp. 1625-1628. [15] E. Alsentzer et al., "Publicly available clinical BERT embeddings," arXiv preprint arXiv:1904.03323, 2019. [16] G. Crichton, S. Pyysalo, B. Chiu, and A. Korhonen, "A neural network multi-task learning approach to biomedical named entity recognition," BMC bioinformatics, vol. 18, no. 1, pp. 1-14, 2017. [17] K. B. Cohen et al., "The colorado richly annotated full text (craft) corpus: Multi-model annotation in the biomedical domain," Handbook of Linguistic annotation, pp. 1379-1394, 2017. [18] C. Vasantharajan, K. Z. Tun, H. Thi-Nga, S. Jain, T. Rong, and C. E. Siong, "MedBERT: A Pre-trained Language Model for Biomedical Named Entity Recognition," 2022: IEEE, doi: 10.23919/apsipaasc55919.2022.9980157. [Online]. Available: https://dx.doi.org/10.23919/apsipaasc55919.2022.9980157 [19] M. Yasunaga et al., "Deep bidirectional language-knowledge graph pretraining," Advances in Neural Information Processing Systems, vol. 35, pp. 37309-37323, 2022. [20] X. Zhang et al., "Greaselm: Graph reasoning enhanced language models for question answering," arXiv preprint arXiv:2201.08860, 2022. [21] Z. Dong and Q. Dong, Hownet and the computation of meaning (with Cd-rom). World Scientific, 2006. [22] B. Xu et al., "CN-DBpedia: A never-ending Chinese knowledge extraction system," in International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems, 2017: Springer, pp. 428-438. [23] A. Saxena, S. Chakrabarti, and P. Talukdar, "Question answering over temporal knowledge graphs," arXiv preprint arXiv:2106.01515, 2021. [24] A. Pal, L. K. Umapathi, and M. Sankarasubbu, "Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering," in Conference on Health, Inference, and Learning, 2022: PMLR, pp. 248-260. [25] D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P. Szolovits, "What disease does this patient have? a large-scale open domain question answering dataset from medical exams," Applied Sciences, vol. 11, no. 14, p. 6421, 2021. [26] K. Degtyarenko et al., "ChEBI: a database and ontology for chemical entities of biological interest," Nucleic acids research, vol. 36, no. suppl_1, pp. D344-D350, 2007. [27] M. Ashburner et al., "Gene ontology: tool for the unification of biology," Nature genetics, vol. 25, no. 1, pp. 25-29, 2000. [28] "NetworkX Home Page." https://networkx.org/ (accessed. [29] K. Singhal et al., "Towards expert-level medical question answering with large language models," arXiv preprint arXiv:2305.09617, 2023. [30] Y. Liu et al., "Roberta: A robustly optimized bert pretraining approach," arXiv preprint arXiv:1907.11692, 2019. [31] C. Raffel et al., "Exploring the limits of transfer learning with a unified text-to-text transformer," Journal of machine learning research, vol. 21, no. 140, pp. 1-67, 2020. [32] E. J. Hu et al., "Lora: Low-rank adaptation of large language models," arXiv preprint arXiv:2106.09685, 2021. |
電子全文 Fulltext |
本電子全文僅授權使用者為學術研究之目的,進行個人非營利性質之檢索、閱讀、列印。請遵守中華民國著作權法之相關規定,切勿任意重製、散佈、改作、轉貼、播送,以免觸法。 論文使用權限 Thesis access permission:自定論文開放時間 user define 開放時間 Available: 校內 Campus:開放下載的時間 available 2027-06-17 校外 Off-campus:開放下載的時間 available 2027-06-17 您的 IP(校外) 位址是 18.97.9.175 現在時間是 2026-08-11 論文校外開放下載的時間是 2027-06-17 Your IP address is 18.97.9.175 The current date is 2026-08-11 This thesis will be available to you on 2027-06-17. |
紙本論文 Printed copies |
紙本論文的公開資訊在102學年度以後相對較為完整。如果需要查詢101學年度以前的紙本論文公開資訊,請聯繫圖資處紙本論文服務櫃台。如有不便之處敬請見諒。 開放時間 available 2027-06-17 |
QR Code |