Responsive image
博碩士論文 etd-0802124-225407 詳細資訊
Title page for etd-0802124-225407
論文名稱
Title
以大型語言模型結構化法院判決書 - 以金融詐欺判決書為例
Structuring Court Judgments with Large Language Models - A Case Study of Financial Fraud Judgments
系所名稱
Department
畢業學年期
Year, semester
語文別
Language
學位類別
Degree
頁數
Number of pages
71
研究生
Author
指導教授
Advisor
召集委員
Convenor
口試委員
Advisory Committee
口試日期
Date of Exam
2024-07-19
繳交日期
Date of Submission
2024-09-02
關鍵字
Keywords
金融詐欺、判決書、結構化數據、自然語言處理、大型語言模型、文本相似度
Financial Fraud, Judgments, Structured Data, Natural Language Processing, Large Language Models, Text Similarity
統計
Statistics
本論文已被瀏覽 572 次,被下載 0
The thesis/dissertation has been browsed 572 times, has been downloaded 0 times.
中文摘要
在當今全球化和數位化的時代,詐欺已成為影響全球經濟和社會結構的嚴重問題。台灣也面臨著日益嚴重的詐欺挑戰,尤其是在數位經濟快速發展的背景下。為了有效預防詐欺犯罪,深入理解犯罪者的行為模式至關重要,而閱讀法律判決書是了解這些資訊的最佳途徑。但台灣的判決書內容繁多且艱澀,對普通民眾甚至專業人士來說理解起來存在困難。
本研究結合了正則表達法與自然語言處理(NLP)技術,針對金融詐欺相關的法院判決書進行結構化處理,並提取出多項關鍵特徵,如違反法條、法院地點、犯罪過程和被告姓名。通過應用Llama 3大型語言模型,本研究克服了文本結構複雜性和格式不統一的挑戰,成功解析了包含犯罪過程等重要資訊的段落。然而,硬體資源的限制和法律專業知識的不足對研究帶來了一定挑戰。我們採用了少量學習(Few-Shot Learning, FSL)技術,並壓縮了模型以適應有限的GPU容量,這些措施雖然提高了處理效率,但可能影響了部分資訊的完整性。未來工作將著重於優化模型的運行效率,並加強與法律專業人士的合作,以期能在法律科技領域做出更具貢獻的研究成果。
Abstract
In the globalized and digital era, fraud has become a critical issue, affecting economies and societies worldwide. Taiwan faces significant fraud challenges, especially amid rapid digital economic growth. Understanding perpetrator behavior through legal court rulings is vital for effective prevention. However, Taiwan's court rulings are often complex and challenging to comprehend.
This study integrates regular expressions with natural language processing (NLP) techniques to structure court rulings on financial fraud, extracting key features such as violated laws, court locations, and criminal processes. Using the Llama 3 large language model, the study addressed challenges in text complexity and inconsistency, though it was limited by hardware constraints and a lack of legal expertise. Few-Shot Learning (FSL) techniques and model compression were employed to enhance processing efficiency, albeit with some potential information loss. Future work will aim to optimize model performance and strengthen collaboration with legal experts to produce more impactful outcomes in legal technology.


目次 Table of Contents
論文審定書 i
致謝 ii
摘要 iii
Abstract iv
Content v
Table of Figures viii
Table of Tables ix
Chapter 1 Introduction 1
1.1 Research Background 1
1.2 Research Motivation & Purpose 6
Chapter 2 Literature Review 8
2.1 Types of Financial Fraud 8
2.2 Definition of Fraud in Taiwan 11
2.3 The Role and Value of Legal Judgments 15
2.4 Applications of Natural Language Processing in Text Analysis 17
2.5 Large Language Models (LLM) 20
Chapter 3 Research Method 23
3.1 Research Framework 23
3.2 Data Collection 26
3.3 Court Judgment Content 27
3.4 Regular Expressions 31
3.5 Natural Language Processing Models (NLP) 34
3.6 Few-Shot Learning (FSL) 34
3.7 Model Selection 38
3.8 Text Similarity Analysis 39
Chapter 4 Experimental Results 42
4.1 Data Preprocessing 42
4.2 Results of Regex Extraction 43
4.3 NLP Extraction Results 46
4.4 Evaluation of Model-Generated Results 52
Chapter 5 Conclusion 55
5.1 Discussions 55
5.2 Limitations and Future Works 57
References 60
參考文獻 References
Anand, V., Tina Dacin, M., & Murphy, P. R. (2015). The continued need for diversity in fraud research. Journal of Business Ethics, 131, 751-755.
Brzozowski, J. A. (1964). Derivatives of regular expressions. Journal of the ACM (JACM), 11(4), 481-494.
Cahill‐O'Callaghan, R. J. (2013). The influence of personal values on legal judgments. Journal of Law and Society, 40(4), 596-623.
Carchiolo, V., Longheu, A., Reitano, G., & Zagarella, L. (2019). Medical prescription classification: a NLP-based approach. 2019 Federated Conference on Computer Science and Information Systems (FedCSIS),
Chen, X., Xie, H., Cheng, G., Poon, L. K., Leng, M., & Wang, F. L. (2020). Trends and features of the applications of natural language processing techniques for clinical trials text analysis. Applied Sciences, 10(6), 2157.
Câmpeanu, C., Salomaa, K., & Yu, S. (2003). A formal study of practical regular expressions. International Journal of Foundations of Computer Science, 14(06), 1007-1018.
Dimmock, S. G., & Gerken, W. C. (2012). Predicting fraud by investment managers. Journal of Financial Economics, 105(1), 153-173.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., & Fan, A. (2024). The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
Hospedales, T., Antoniou, A., Micaelli, P., & Storkey, A. (2021). Meta-learning in neural networks: A survey. IEEE transactions on pattern analysis and machine intelligence, 44(9), 5149-5169.
Kulis, B. (2013). Metric learning: A survey. Foundations and Trends® in Machine Learning, 5(4), 287-364.
Lenci, A., Montemagni, S., Pirrelli, V., & Venturi, G. (2007). NLP-based ontology learning from legal texts. A case study. LOAIT, 321, 113-129.
Li, Y., Pan, Q., Wang, S., Yang, T., & Cambria, E. (2018). A generative model for category text generation. Information Sciences, 450, 301-315.
Lim, A. (2022). Exploring Dating Apps: Catfishing or Kittenfishing? University of Akron].
Merchant, K., & Pande, Y. (2018). Nlp based latent semantic analysis for legal text summarization. 2018 international conference on advances in computing, communications and informatics (ICACCI),
Ozili, P. K. (2020). Advances and issues in fraud research: a commentary. Journal of Financial Crime, 27(1), 92-103.
Shahmirzadi, O., Lugowski, A., & Younge, K. (2019). Text similarity in vector space models: a comparative study. 2019 18th IEEE international conference on machine learning and applications (ICMLA),
Singh, R., & Singh, S. (2021). Text similarity measures in news articles by vector space model using NLP. Journal of The Institution of Engineers (India): Series B, 102, 329-338.
Sithic, H. L., & Balasubramanian, T. (2013). Survey of insurance fraud detection using data mining techniques. arXiv preprint arXiv:1309.0806.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., & Azhar, F. (2023). Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
Villa, V. (1997). Legal theory and value judgments. Law and Philosophy, 16, 447-477.
Wang, Y., Yao, Q., Kwok, J. T., & Ni, L. M. (2020). Generalizing from a few examples: A survey on few-shot learning. ACM computing surveys (csur), 53(3), 1-34.
Xia, P., Zhang, L., & Li, F. (2015). Learning similarity with cosine similarity ensemble. Information Sciences, 307, 39-52.
Xing, F. Z., Cambria, E., & Welsch, R. E. (2018). Natural language based financial forecasting: a survey. Artificial Intelligence Review, 50(1), 49-73.
Zhong, H., Xiao, C., Tu, C., Zhang, T., Liu, Z., & Sun, M. (2020). How does NLP benefit legal system: A summary of legal artificial intelligence. arXiv preprint arXiv:2004.12158.
Zojaji, Z., Atani, R. E., & Monadjemi, A. H. (2016). A survey of credit card fraud detection techniques: data and technique oriented perspective. arXiv preprint arXiv:1611.06439.
黃資閔. (2024). 詐欺案件之人頭帳戶角色與行為分析;基於司法判決的實證研究 國立中山大學]. 臺灣博碩士論文知識加值系統. 高雄市. https://hdl.handle.net/11296/tg6cbq
電子全文 Fulltext
本電子全文僅授權使用者為學術研究之目的,進行個人非營利性質之檢索、閱讀、列印。請遵守中華民國著作權法之相關規定,切勿任意重製、散佈、改作、轉貼、播送,以免觸法。
論文使用權限 Thesis access permission:自定論文開放時間 user define
開放時間 Available:
校內 Campus:開放下載的時間 available 2034-09-02
校外 Off-campus:開放下載的時間 available 2034-09-02

您的 IP(校外) 位址是 18.97.14.91
現在時間是 2026-07-15
論文校外開放下載的時間是 2034-09-02

Your IP address is 18.97.14.91
The current date is 2026-07-15
This thesis will be available to you on 2034-09-02.

紙本論文 Printed copies
紙本論文的公開資訊在102學年度以後相對較為完整。如果需要查詢101學年度以前的紙本論文公開資訊,請聯繫圖資處紙本論文服務櫃台。如有不便之處敬請見諒。
開放時間 available 2029-09-02

QR Code