論文使用權限 Thesis access permission:校內校外完全公開 unrestricted
開放時間 Available:
校內 Campus: 已公開 available
校外 Off-campus: 已公開 available
論文名稱 Title |
最佳化具延遲效應的藥物治療決策--使用時序動態決策轉換器 Temporal-Dynamic Decision Transformer for Optimizing Drug Therapy with Delayed Effects |
||
系所名稱 Department |
|||
畢業學年期 Year, semester |
語文別 Language |
||
學位類別 Degree |
頁數 Number of pages |
58 |
|
研究生 Author |
|||
指導教授 Advisor |
|||
召集委員 Convenor |
|||
口試委員 Advisory Committee |
|||
口試日期 Date of Exam |
2024-07-11 |
繳交日期 Date of Submission |
2024-08-21 |
關鍵字 Keywords |
強化學習、決策轉換器、序列建模、藥物延遲效應、時間變化資料 Reinforcement Learning, Decision Transformer, Sequence Modeling, Delayed Effects of Drugs, Temporal Data |
||
統計 Statistics |
本論文已被瀏覽 389 次,被下載 5 次 The thesis/dissertation has been browsed 389 times, has been downloaded 5 times. |
中文摘要 |
隨著科技進步及資訊爆炸,機器學習模型越來越多地被用於輔助藥物治療決策。然而,隨著這些模型的廣泛應用,我們面臨著一個重要挑戰:藥物效應的延遲性。傳統的機器學習模型通常難以有效處理藥物在人體內隨時間變化的複雜效應,這可能導致治療效果的誤判並選擇次優的治療方案。 在現有的機器學習模型優化藥物開立狀況中,大多著重於短期效果的預測和即時調整,忽視了藥物長期作用的累積效應和生理參數的穩定性。這種忽視可能導致治療過程中的波動,甚至對患者造成潛在的健康風險。為了解決這個問題,需要一種能夠全面考慮藥物延遲效應和整個治療軌跡的新方法。 本論文通過引入軌跡感知優化和考慮藥物延遲效應的複雜獎勵結構,有效地將整個治療過程納入考量,旨在優化具有延遲效應的藥物療法。這種方法不僅關注最終治療目標,還確保在整個治療過程中維持穩定的生理參數。 |
Abstract |
With technological advancements and the explosion of information, machine learning models are increasingly being used to assist in drug therapy decision-making. However, as these models become more widely applied, we face a significant challenge: the delayed effects of medications. Traditional machine learning models often struggle to effectively handle the complex temporal dynamics of drugs in the human body, which can lead to misestimation of treatment efficacy and suboptimal therapeutic regimens. In existing machine learning models optimizing drug prescription scenarios, the focus is primarily on predicting short-term effects and making immediate adjustments, neglecting the cumulative impact of long-term drug actions and the stability of physiological parameters. This oversight can result in fluctuations during the treatment process, potentially posing health risks to patients. To address this issue, a new approach is needed that comprehensively considers both the delayed effects of drugs and the entire treatment trajectory. This paper introduces an innovative reinforcement learning method that incorporates trajectory-aware optimization and a sophisticated reward structure accounting for delayed drug effects. This approach effectively considers the entire treatment process, aiming to optimize drug therapies with delayed effects. The method not only focuses on the final treatment goal but also ensures the maintenance of stable physiological parameters throughout the entire treatment process. |
目次 Table of Contents |
論文審定書 i 摘要 ii Abstract iii Table of Figures v Table of Tables vi 1. Introduction 1 2. Background 3 2.1 Markov Decision Process 3 2.2 Reinforcement Learning 5 2.2.1 Types of Reinforcement Learning 7 2.2.2 Reward Design and Strategies 10 2.2.3 Algorithms in Reinforcement Learning 12 2.3 Sequential Data 15 2.4 Transformers 16 2.5 Decision Transformer 17 2.6 Delay Effect 19 3. Methodology 21 4. Experiment Results 31 4.1 Data Source 31 4.2 Measurement 33 4.2.1 Status classification 33 4.2.2 Index evaluation 36 4.3 Experiment Results 38 4.3.1 Comparison of Models 38 4.3.2 Impact of Custom Loss Function 42 5. Conclusion 43 References 45 |
參考文獻 References |
Ba, J. L., Kiros, J. R., & Hinton, G. E. (2016). Layer Normalization (arXiv:1607.06450). arXiv. https://doi.org/10.48550/arXiv.1607.06450 Bellman, R. (1957). A Markovian Decision Process. Indiana University Mathematics Journal, 6(4), 679–684. https://doi.org/10.1512/iumj.1957.6.56038 Bennett, C. C., & Hauser, K. (2013). Artificial intelligence framework for simulating clinical decision-making: A Markov decision process approach. Artificial Intelligence in Medicine, 57(1), 9–19. https://doi.org/10.1016/j.artmed.2012.12.003 Besarab, A., Bolton, W. K., Browne, J. K., Egrie, J. C., Nissenson, A. R., Okamoto, D. M., Schwab, S. J., & Goodkin, D. A. (1998). The Effects of Normal as Compared with Low Hematocrit Values in Patients with Cardiac Disease Who Are Receiving Hemodialysis and Epoetin. New England Journal of Medicine, 339(9), 584–590. https://doi.org/10.1056/NEJM199808273390903 Chapelle, O., & Li, L. (2011). An Empirical Evaluation of Thompson Sampling. Advances in Neural Information Processing Systems, 24. https://papers.nips.cc/paper_files/paper/2011/hash/e53a0a2978c28872a4505bdb51db06dc-Abstract.html Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., & Mordatch, I. (2021). Decision Transformer: Reinforcement Learning via Sequence Modeling (arXiv:2106.01345). arXiv. http://arxiv.org/abs/2106.01345 Chiu, Y.-W., Lin, M.-Y., Yen, H.-R., Hsu, C., Ku, C.-T., & Kang, Y. (2023). Using an Ensemble Model to Improve ESA Prescription in Hemodialysis (HD): TH-PO042. Journal of the American Society of Nephrology, 34(11S), 101. https://doi.org/10.1681/ASN.20233411S1101a Choi, E., Bahadori, M. T., Schuetz, A., Stewart, W. F., & Sun, J. (2016). Doctor AI: Predicting Clinical Events via Recurrent Neural Networks (arXiv:1511.05942). arXiv. https://doi.org/10.48550/arXiv.1511.05942 Deffayet, R., Thonet, T., Renders, J.-M., & de Rijke, M. (2023). Offline Evaluation for Reinforcement Learning-Based Recommendation: A Critical Issue and Some Alternatives. SIGIR Forum, 56(2), 3:1-3:14. https://doi.org/10.1145/3582900.3582905 Del Vecchio, L., & Minutolo, R. (2021). ESA, Iron Therapy and New Drugs: Are There New Perspectives in the Treatment of Anaemia? Journal of Clinical Medicine, 10(4), 839. https://doi.org/10.3390/jcm10040839 Eschbach, J. W. (2002). Anemia management in chronic kidney disease: Role of factors affecting epoetin responsiveness. Journal of the American Society of Nephrology: JASN, 13(5), 1412–1414. https://doi.org/10.1097/01.asn.0000016440.52271.f7 Eschmann, J. (2021). Reward Function Design in Reinforcement Learning. In B. Belousov, H. Abdulsamad, P. Klink, S. Parisi, & J. Peters (Eds.), Reinforcement Learning Algorithms: Analysis and Applications (pp. 25–33). Springer International Publishing. https://doi.org/10.1007/978-3-030-41188-6_3 Fatemi, M., Wu, M., Petch, J., Nelson, W., Connolly, S. J., Benz, A., Carnicelli, A., & Ghassemi, M. (2022). Semi-Markov Offline Reinforcement Learning for Healthcare. Proceedings of the Conference on Health, Inference, and Learning, 119–137. https://proceedings.mlr.press/v174/fatemi22a.html Fokkema, M., Smits, N., Zeileis, A., Hothorn, T., & Kelderman, H. (2018). Detecting treatment-subgroup interactions in clustered data with generalized linear mixed-effects model trees. Behavior Research Methods, 50(5), 2016–2034. https://doi.org/10.3758/s13428-017-0971-x Gaweda, A. E., Muezzinoglu, M. K., Aronoff, G. R., Jacobs, A. A., Zurada, J. M., & Brier, M. E. (2005). Individualization of pharmacological anemia management using reinforcement learning. Neural Networks, 18(5–6), 826–834. https://doi.org/10.1016/j.neunet.2005.06.020 Ghahramani, Z. (2001). An introduction to hidden markov models and bayesian networks. International Journal of Pattern Recognition and Artificial Intelligence, 15(01), 9–42. https://doi.org/10.1142/S0218001401000836 Hanafusa, N., Nakai, S., Iseki, K., & Tsubakihara, Y. (2015). Japanese society for dialysis therapy renal data registry—A window through which we can view the details of Japanese dialysis population. Kidney International Supplements, 5(1), 15–22. https://doi.org/10.1038/kisup.2015.5 Hochreiter, S., & Schmidhuber, J. (1997). Long Short-Term Memory. Neural Computation, 9(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735 Kaelbling, L. P., Littman, M. L., & Moore, A. W. (1996). Reinforcement Learning: A Survey (arXiv:cs/9605103). arXiv. https://doi.org/10.48550/arXiv.cs/9605103 Kumar, A., Zhou, A., Tucker, G., & Levine, S. (2020). Conservative Q-Learning for Offline Reinforcement Learning. Advances in Neural Information Processing Systems, 33, 1179–1191. https://proceedings.neurips.cc/paper/2020/hash/0d2b2061826a5df3221116a5085a6052-Abstract.html Längkvist, M., Karlsson, L., & Loutfi, A. (2014). A review of unsupervised feature learning and deep learning for time-series modeling. Pattern Recognition Letters, 42, 11–24. https://doi.org/10.1016/j.patrec.2014.01.008 Levine, S., Kumar, A., Tucker, G., & Fu, J. (2020). Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems (arXiv:2005.01643). arXiv. https://doi.org/10.48550/arXiv.2005.01643 Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., & Wierstra, D. (2019). Continuous control with deep reinforcement learning (arXiv:1509.02971). arXiv. http://arxiv.org/abs/1509.02971 Lipton, Z. C., Kale, D. C., & Wetzel, R. (2016). Modeling Missing Data in Clinical Time Series with RNNs (arXiv:1606.04130). arXiv. https://doi.org/10.48550/arXiv.1606.04130 Machado-Vieira, R., Salvadore, G., Luckenbaugh, D. A., Manji, H. K., & Zarate, C. A. (2008). Rapid onset of antidepressant action: A new paradigm in the research and treatment of major depressive disorder. The Journal of Clinical Psychiatry, 69(6), 946–958. https://doi.org/10.4088/jcp.v69n0610 Malof, J. M., & Gaweda, A. E. (2011). Optimizing drug therapy with Reinforcement Learning: The case of Anemia Management. The 2011 International Joint Conference on Neural Networks, 2088–2092. https://doi.org/10.1109/IJCNN.2011.6033485 Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533. https://doi.org/10.1038/nature14236 Moerland, T. M., Broekens, J., Plaat, A., & Jonker, C. M. (2022). Model-based Reinforcement Learning: A Survey (arXiv:2006.16712). arXiv. https://doi.org/10.48550/arXiv.2006.16712 Nakhoul, G., & Simon, J. F. (2016). Anemia of chronic kidney disease: Treat it, but not too aggressively. Cleveland Clinic Journal of Medicine, 83(8), 613–624. https://doi.org/10.3949/ccjm.83a.15065 Nemati, S., Ghassemi, M. M., & Clifford, G. D. (2016). Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach. Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference, 2016, 2978–2981. https://doi.org/10.1109/EMBC.2016.7591355 Nguyen, H. T., & Luong, N. H. (2021). Applying Deep Reinforcement Learning in Automated Stock Trading. In N. H. Phuong & V. Kreinovich (Eds.), Soft Computing: Biomedical and Related Applications (pp. 285–297). Springer International Publishing. https://doi.org/10.1007/978-3-030-76620-7_25 Off-Policy Deep Reinforcement Learning without Exploration. (n.d.). Retrieved June 28, 2024, from https://proceedings.mlr.press/v97/fujimoto19a.html Puterman, M. L. (2005). Markov Decision Processes: Discrete Stochastic Dynamic Programming (1st ed.). Wiley-Interscience. Random Forests | Machine Language. (n.d.). Retrieved July 5, 2024, from https://dl.acm.org/doi/abs/10.1023/A:1010933404324 Robins, J. M., Hernán, M. A., & Brumback, B. (2000). Marginal structural models and causal inference in epidemiology. Epidemiology (Cambridge, Mass.), 11(5), 550–560. https://doi.org/10.1097/00001648-200009000-00011 Schulman, J., Levine, S., Moritz, P., Jordan, M. I., & Abbeel, P. (2017). Trust Region Policy Optimization (arXiv:1502.05477). arXiv. https://doi.org/10.48550/arXiv.1502.05477 Sutton, R. S., & Barto, A. G. (n.d.). Reinforcement Learning: An Introduction. Time Series Analysis: Forecasting and Control, 5th Edition | Wiley. (n.d.). Retrieved June 29, 2024, from https://www.wiley.com/en-us/Time+Series+Analysis%3A+Forecasting+and+Control%2C+5th+Edition-p-9781118675021 Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2023). Attention Is All You Need (arXiv:1706.03762). arXiv. https://doi.org/10.48550/arXiv.1706.03762 Veviurko, G., Böhmer, W., & de Weerdt, M. (2024). To the Max: Reinventing Reward in Reinforcement Learning (arXiv:2402.01361). arXiv. https://doi.org/10.48550/arXiv.2402.01361 Wagenmaker, A., & Pacchiano, A. (2023). Leveraging Offline Data in Online Reinforcement Learning (arXiv:2211.04974). arXiv. http://arxiv.org/abs/2211.04974 Waring, R. (2006). Correction of Anemia with Epoetin Alfa in Chronic Kidney Disease. The New England Journal of Medicine. Zhao, L., Hu, C., Cheng, J., Zhang, P., Jiang, H., & Chen, J. (2019). Haemoglobin variability and all‐cause mortality in haemodialysis patients: A systematic review and meta‐analysis. Nephrology, 24(12), 1265–1272. https://doi.org/10.1111/nep.13560 |
電子全文 Fulltext |
本電子全文僅授權使用者為學術研究之目的,進行個人非營利性質之檢索、閱讀、列印。請遵守中華民國著作權法之相關規定,切勿任意重製、散佈、改作、轉貼、播送,以免觸法。 論文使用權限 Thesis access permission:校內校外完全公開 unrestricted 開放時間 Available: 校內 Campus: 已公開 available 校外 Off-campus: 已公開 available |
紙本論文 Printed copies |
紙本論文的公開資訊在102學年度以後相對較為完整。如果需要查詢101學年度以前的紙本論文公開資訊,請聯繫圖資處紙本論文服務櫃台。如有不便之處敬請見諒。 開放時間 available 已公開 available |
QR Code |