Responsive image
博碩士論文 etd-0721124-024359 詳細資訊
Title page for etd-0721124-024359
論文名稱
Title
最佳化具延遲效應的藥物治療決策--使用時序動態決策轉換器
Temporal-Dynamic Decision Transformer for Optimizing Drug Therapy with Delayed Effects
系所名稱
Department
畢業學年期
Year, semester
語文別
Language
學位類別
Degree
頁數
Number of pages
58
研究生
Author
指導教授
Advisor
召集委員
Convenor
口試委員
Advisory Committee
口試日期
Date of Exam
2024-07-11
繳交日期
Date of Submission
2024-08-21
關鍵字
Keywords
強化學習、決策轉換器、序列建模、藥物延遲效應、時間變化資料
Reinforcement Learning, Decision Transformer, Sequence Modeling, Delayed Effects of Drugs, Temporal Data
統計
Statistics
本論文已被瀏覽 389 次,被下載 5
The thesis/dissertation has been browsed 389 times, has been downloaded 5 times.
中文摘要
隨著科技進步及資訊爆炸,機器學習模型越來越多地被用於輔助藥物治療決策。然而,隨著這些模型的廣泛應用,我們面臨著一個重要挑戰:藥物效應的延遲性。傳統的機器學習模型通常難以有效處理藥物在人體內隨時間變化的複雜效應,這可能導致治療效果的誤判並選擇次優的治療方案。
在現有的機器學習模型優化藥物開立狀況中,大多著重於短期效果的預測和即時調整,忽視了藥物長期作用的累積效應和生理參數的穩定性。這種忽視可能導致治療過程中的波動,甚至對患者造成潛在的健康風險。為了解決這個問題,需要一種能夠全面考慮藥物延遲效應和整個治療軌跡的新方法。
本論文通過引入軌跡感知優化和考慮藥物延遲效應的複雜獎勵結構,有效地將整個治療過程納入考量,旨在優化具有延遲效應的藥物療法。這種方法不僅關注最終治療目標,還確保在整個治療過程中維持穩定的生理參數。
Abstract
With technological advancements and the explosion of information, machine learning models are increasingly being used to assist in drug therapy decision-making. However, as these models become more widely applied, we face a significant challenge: the delayed effects of medications. Traditional machine learning models often struggle to effectively handle the complex temporal dynamics of drugs in the human body, which can lead to misestimation of treatment efficacy and suboptimal therapeutic regimens.
In existing machine learning models optimizing drug prescription scenarios, the focus is primarily on predicting short-term effects and making immediate adjustments, neglecting the cumulative impact of long-term drug actions and the stability of physiological parameters. This oversight can result in fluctuations during the treatment process, potentially posing health risks to patients. To address this issue, a new approach is needed that comprehensively considers both the delayed effects of drugs and the entire treatment trajectory.
This paper introduces an innovative reinforcement learning method that incorporates trajectory-aware optimization and a sophisticated reward structure accounting for delayed drug effects. This approach effectively considers the entire treatment process, aiming to optimize drug therapies with delayed effects. The method not only focuses on the final treatment goal but also ensures the maintenance of stable physiological parameters throughout the entire treatment process.
目次 Table of Contents
論文審定書 i
摘要 ii
Abstract iii
Table of Figures v
Table of Tables vi
1. Introduction 1
2. Background 3
2.1 Markov Decision Process 3
2.2 Reinforcement Learning 5
2.2.1 Types of Reinforcement Learning 7
2.2.2 Reward Design and Strategies 10
2.2.3 Algorithms in Reinforcement Learning 12
2.3 Sequential Data 15
2.4 Transformers 16
2.5 Decision Transformer 17
2.6 Delay Effect 19
3. Methodology 21
4. Experiment Results 31
4.1 Data Source 31
4.2 Measurement 33
4.2.1 Status classification 33
4.2.2 Index evaluation 36
4.3 Experiment Results 38
4.3.1 Comparison of Models 38
4.3.2 Impact of Custom Loss Function 42
5. Conclusion 43
References 45

參考文獻 References
Ba, J. L., Kiros, J. R., & Hinton, G. E. (2016). Layer Normalization (arXiv:1607.06450). arXiv. https://doi.org/10.48550/arXiv.1607.06450
Bellman, R. (1957). A Markovian Decision Process. Indiana University Mathematics Journal, 6(4), 679–684. https://doi.org/10.1512/iumj.1957.6.56038
Bennett, C. C., & Hauser, K. (2013). Artificial intelligence framework for simulating clinical decision-making: A Markov decision process approach. Artificial Intelligence in Medicine, 57(1), 9–19. https://doi.org/10.1016/j.artmed.2012.12.003
Besarab, A., Bolton, W. K., Browne, J. K., Egrie, J. C., Nissenson, A. R., Okamoto, D. M., Schwab, S. J., & Goodkin, D. A. (1998). The Effects of Normal as Compared with Low Hematocrit Values in Patients with Cardiac Disease Who Are Receiving Hemodialysis and Epoetin. New England Journal of Medicine, 339(9), 584–590. https://doi.org/10.1056/NEJM199808273390903
Chapelle, O., & Li, L. (2011). An Empirical Evaluation of Thompson Sampling. Advances in Neural Information Processing Systems, 24. https://papers.nips.cc/paper_files/paper/2011/hash/e53a0a2978c28872a4505bdb51db06dc-Abstract.html
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., & Mordatch, I. (2021). Decision Transformer: Reinforcement Learning via Sequence Modeling (arXiv:2106.01345). arXiv. http://arxiv.org/abs/2106.01345
Chiu, Y.-W., Lin, M.-Y., Yen, H.-R., Hsu, C., Ku, C.-T., & Kang, Y. (2023). Using an Ensemble Model to Improve ESA Prescription in Hemodialysis (HD): TH-PO042. Journal of the American Society of Nephrology, 34(11S), 101. https://doi.org/10.1681/ASN.20233411S1101a
Choi, E., Bahadori, M. T., Schuetz, A., Stewart, W. F., & Sun, J. (2016). Doctor AI: Predicting Clinical Events via Recurrent Neural Networks (arXiv:1511.05942). arXiv. https://doi.org/10.48550/arXiv.1511.05942
Deffayet, R., Thonet, T., Renders, J.-M., & de Rijke, M. (2023). Offline Evaluation for Reinforcement Learning-Based Recommendation: A Critical Issue and Some Alternatives. SIGIR Forum, 56(2), 3:1-3:14. https://doi.org/10.1145/3582900.3582905
Del Vecchio, L., & Minutolo, R. (2021). ESA, Iron Therapy and New Drugs: Are There New Perspectives in the Treatment of Anaemia? Journal of Clinical Medicine, 10(4), 839. https://doi.org/10.3390/jcm10040839
Eschbach, J. W. (2002). Anemia management in chronic kidney disease: Role of factors affecting epoetin responsiveness. Journal of the American Society of Nephrology: JASN, 13(5), 1412–1414. https://doi.org/10.1097/01.asn.0000016440.52271.f7
Eschmann, J. (2021). Reward Function Design in Reinforcement Learning. In B. Belousov, H. Abdulsamad, P. Klink, S. Parisi, & J. Peters (Eds.), Reinforcement Learning Algorithms: Analysis and Applications (pp. 25–33). Springer International Publishing. https://doi.org/10.1007/978-3-030-41188-6_3
Fatemi, M., Wu, M., Petch, J., Nelson, W., Connolly, S. J., Benz, A., Carnicelli, A., & Ghassemi, M. (2022). Semi-Markov Offline Reinforcement Learning for Healthcare. Proceedings of the Conference on Health, Inference, and Learning, 119–137. https://proceedings.mlr.press/v174/fatemi22a.html
Fokkema, M., Smits, N., Zeileis, A., Hothorn, T., & Kelderman, H. (2018). Detecting treatment-subgroup interactions in clustered data with generalized linear mixed-effects model trees. Behavior Research Methods, 50(5), 2016–2034. https://doi.org/10.3758/s13428-017-0971-x
Gaweda, A. E., Muezzinoglu, M. K., Aronoff, G. R., Jacobs, A. A., Zurada, J. M., & Brier, M. E. (2005). Individualization of pharmacological anemia management using reinforcement learning. Neural Networks, 18(5–6), 826–834. https://doi.org/10.1016/j.neunet.2005.06.020
Ghahramani, Z. (2001). An introduction to hidden markov models and bayesian networks. International Journal of Pattern Recognition and Artificial Intelligence, 15(01), 9–42. https://doi.org/10.1142/S0218001401000836
Hanafusa, N., Nakai, S., Iseki, K., & Tsubakihara, Y. (2015). Japanese society for dialysis therapy renal data registry—A window through which we can view the details of Japanese dialysis population. Kidney International Supplements, 5(1), 15–22. https://doi.org/10.1038/kisup.2015.5
Hochreiter, S., & Schmidhuber, J. (1997). Long Short-Term Memory. Neural Computation, 9(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735
Kaelbling, L. P., Littman, M. L., & Moore, A. W. (1996). Reinforcement Learning: A Survey (arXiv:cs/9605103). arXiv. https://doi.org/10.48550/arXiv.cs/9605103
Kumar, A., Zhou, A., Tucker, G., & Levine, S. (2020). Conservative Q-Learning for Offline Reinforcement Learning. Advances in Neural Information Processing Systems, 33, 1179–1191. https://proceedings.neurips.cc/paper/2020/hash/0d2b2061826a5df3221116a5085a6052-Abstract.html
Längkvist, M., Karlsson, L., & Loutfi, A. (2014). A review of unsupervised feature learning and deep learning for time-series modeling. Pattern Recognition Letters, 42, 11–24. https://doi.org/10.1016/j.patrec.2014.01.008
Levine, S., Kumar, A., Tucker, G., & Fu, J. (2020). Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems (arXiv:2005.01643). arXiv. https://doi.org/10.48550/arXiv.2005.01643
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., & Wierstra, D. (2019). Continuous control with deep reinforcement learning (arXiv:1509.02971). arXiv. http://arxiv.org/abs/1509.02971
Lipton, Z. C., Kale, D. C., & Wetzel, R. (2016). Modeling Missing Data in Clinical Time Series with RNNs (arXiv:1606.04130). arXiv. https://doi.org/10.48550/arXiv.1606.04130
Machado-Vieira, R., Salvadore, G., Luckenbaugh, D. A., Manji, H. K., & Zarate, C. A. (2008). Rapid onset of antidepressant action: A new paradigm in the research and treatment of major depressive disorder. The Journal of Clinical Psychiatry, 69(6), 946–958. https://doi.org/10.4088/jcp.v69n0610
Malof, J. M., & Gaweda, A. E. (2011). Optimizing drug therapy with Reinforcement Learning: The case of Anemia Management. The 2011 International Joint Conference on Neural Networks, 2088–2092. https://doi.org/10.1109/IJCNN.2011.6033485
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533. https://doi.org/10.1038/nature14236
Moerland, T. M., Broekens, J., Plaat, A., & Jonker, C. M. (2022). Model-based Reinforcement Learning: A Survey (arXiv:2006.16712). arXiv. https://doi.org/10.48550/arXiv.2006.16712
Nakhoul, G., & Simon, J. F. (2016). Anemia of chronic kidney disease: Treat it, but not too aggressively. Cleveland Clinic Journal of Medicine, 83(8), 613–624. https://doi.org/10.3949/ccjm.83a.15065
Nemati, S., Ghassemi, M. M., & Clifford, G. D. (2016). Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach. Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference, 2016, 2978–2981. https://doi.org/10.1109/EMBC.2016.7591355
Nguyen, H. T., & Luong, N. H. (2021). Applying Deep Reinforcement Learning in Automated Stock Trading. In N. H. Phuong & V. Kreinovich (Eds.), Soft Computing: Biomedical and Related Applications (pp. 285–297). Springer International Publishing. https://doi.org/10.1007/978-3-030-76620-7_25
Off-Policy Deep Reinforcement Learning without Exploration. (n.d.). Retrieved June 28, 2024, from https://proceedings.mlr.press/v97/fujimoto19a.html
Puterman, M. L. (2005). Markov Decision Processes: Discrete Stochastic Dynamic Programming (1st ed.). Wiley-Interscience.
Random Forests | Machine Language. (n.d.). Retrieved July 5, 2024, from https://dl.acm.org/doi/abs/10.1023/A:1010933404324
Robins, J. M., Hernán, M. A., & Brumback, B. (2000). Marginal structural models and causal inference in epidemiology. Epidemiology (Cambridge, Mass.), 11(5), 550–560. https://doi.org/10.1097/00001648-200009000-00011
Schulman, J., Levine, S., Moritz, P., Jordan, M. I., & Abbeel, P. (2017). Trust Region Policy Optimization (arXiv:1502.05477). arXiv. https://doi.org/10.48550/arXiv.1502.05477
Sutton, R. S., & Barto, A. G. (n.d.). Reinforcement Learning: An Introduction.
Time Series Analysis: Forecasting and Control, 5th Edition | Wiley. (n.d.). Retrieved June 29, 2024, from https://www.wiley.com/en-us/Time+Series+Analysis%3A+Forecasting+and+Control%2C+5th+Edition-p-9781118675021
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2023). Attention Is All You Need (arXiv:1706.03762). arXiv. https://doi.org/10.48550/arXiv.1706.03762
Veviurko, G., Böhmer, W., & de Weerdt, M. (2024). To the Max: Reinventing Reward in Reinforcement Learning (arXiv:2402.01361). arXiv. https://doi.org/10.48550/arXiv.2402.01361
Wagenmaker, A., & Pacchiano, A. (2023). Leveraging Offline Data in Online Reinforcement Learning (arXiv:2211.04974). arXiv. http://arxiv.org/abs/2211.04974
Waring, R. (2006). Correction of Anemia with Epoetin Alfa in Chronic Kidney Disease. The New England Journal of Medicine.
Zhao, L., Hu, C., Cheng, J., Zhang, P., Jiang, H., & Chen, J. (2019). Haemoglobin variability and all‐cause mortality in haemodialysis patients: A systematic review and meta‐analysis. Nephrology, 24(12), 1265–1272. https://doi.org/10.1111/nep.13560
電子全文 Fulltext
本電子全文僅授權使用者為學術研究之目的,進行個人非營利性質之檢索、閱讀、列印。請遵守中華民國著作權法之相關規定,切勿任意重製、散佈、改作、轉貼、播送,以免觸法。
論文使用權限 Thesis access permission:校內校外完全公開 unrestricted
開放時間 Available:
校內 Campus: 已公開 available
校外 Off-campus: 已公開 available


紙本論文 Printed copies
紙本論文的公開資訊在102學年度以後相對較為完整。如果需要查詢101學年度以前的紙本論文公開資訊,請聯繫圖資處紙本論文服務櫃台。如有不便之處敬請見諒。
開放時間 available 已公開 available

QR Code