Responsive image
博碩士論文 etd-0708124-155507 詳細資訊
Title page for etd-0708124-155507
論文名稱
Title
多模態迷因檢索與優化之研究
Research on Multimodal Meme Retrieval and Optimization
系所名稱
Department
畢業學年期
Year, semester
語文別
Language
學位類別
Degree
頁數
Number of pages
48
研究生
Author
指導教授
Advisor
召集委員
Convenor
口試委員
Advisory Committee
口試日期
Date of Exam
2024-07-31
繳交日期
Date of Submission
2024-08-08
關鍵字
Keywords
多模態、資訊檢索、社群媒體、隱喻、迷因理解
Multimodal, Information Retrieval, Social Media, Metaphor, Meme Understanding
統計
Statistics
本論文已被瀏覽 524 次,被下載 0
The thesis/dissertation has been browsed 524 times, has been downloaded 0 times.
中文摘要
隨著電腦與網路技術的快速發展,使用者生成內容(User-Generated Content)的儲存成本與傳播所需時間大幅降低,也因此它在短短幾年內以指數性快速增加。根據統計顯示,多媒體內容(Multimedia)讓用戶在網路上進行最多互動與資訊交流的一種類型,其包含圖像(Image)、文字(Text)、影音(Video)等。而隨著多媒體內容的大量產生,許多對於多媒體內容的研究,如內容生成(Content Generation)、資訊檢索(Information Retrieval)等也被陸續提出。
在各種形式的多媒體內容中,迷因(Meme)是一種較為特別的呈現型態,它能以簡短的語句結合圖片傳達個人思想、風格、行為等,它也可能蘊含著文化、時空特定現象等,也因此近年來迷因在全球的社群媒體被廣泛使用。也有越來越多與迷因理解相關的研究問世,如迷因分類、迷因生成等,但卻幾乎沒有迷因檢索相關的研究,這可能是因為對迷因資料通常包含隱喻(Metaphor),同樣的圖片搭配上不同的語句可能就是意義截然不同的內容,加上迷因普遍是由個人所產生,根據主觀性的不同也會影響傳遞的語意,因此要如何有效的檢索迷因是個具挑戰性的任務。本研究提出以多模態方法應用在迷因檢索任務上,並嘗試使用迷因生成模型輔助檢索任務,提升模型檢索表現。
在迷因檢索的績效評估方面,本研究使用Recall@k作為模型的評估指標。研究結果顯示,多模態方法以及迷因生成都能有效應用在迷因檢索任務上,並且也能提升迷因檢索任務的表現。
Abstract
With the rapid advancement of computer and internet technology, the cost of storing and sharing user-generated content has significantly decreased, leading to an exponential increase in such content. Multimedia content, including images, text, and videos, is the most interacted with and exchanged online. Consequently, research on multimedia topics, like content generation and information retrieval, has emerged.
Memes, a unique form of multimedia, combine images with brief text to convey personal thoughts, styles, and cultural phenomena, making them widely popular on social media. Although studies on meme classification and generation are increasing, research on meme retrieval remains scarce due to the subjective and metaphorical nature of memes. This study proposes a multimodal approach to meme retrieval, utilizing a meme generation model to enhance search accuracy.
The study evaluates meme retrieval performance using Recall, demonstrating that the multimodal approach and meme generation effectively improve retrieval outcomes.
目次 Table of Contents
論文審定書 i
誌謝 ii
摘要 iii
Abstract iv
目錄 v
圖次 vii
表次 viii
第一章 緒論 1
1.1 研究背景 1
1.2 研究動機 1
1.3 研究目的 2
第二章 文獻探討 3
2.1 迷因(Meme) 3
2.2 多模態(Multimodal) 3
2.3 資訊檢索(Information Retrieval) 4
2.3.1 文字檢索(Text Retrieval) 4
2.3.2 圖像檢索(Image Retrieval) 4
2.3.3 多模態檢索(Multimodal Retrieval) 5
2.3.3 迷因搜尋引擎(Multimodal Search Engine) 5
2.4 光學字元辨識(Optical Character Recognition) 6
2.5 語意搜尋(Semantic Search) 6
2.6 多模態融合(Multimodal Fusion) 7
2.6.1 早期融合(Early Fusion) 7
2.6.2 晚期融合(Late Fusion) 7
2.7 單模態/多模態模型(Unimodal/Multimodal Model) 8
2.7.1 BERT(Bidirectional Encoder Representations for Transformers) 8
2.7.2卷積神經網路(CNN, Convolutional Neural Network) 9
2.7.3 ViLBERT(Vision-and-Language BERT) 9
2.8 迷因分類(Meme Classification) 10
2.9 迷因生成(Meme Generation) 11
第三章 研究流程 12
3.1 迷因檢索模型(Meme Retrieval Model) 13
3.2 迷因生成模型(Meme Generation Model) 14
3.3 相似度計算(Similarity Calculation) 15
3.4 實際流程 15
第四章 研究結果 17
4.1 資料集介紹 17
4.1.1 MemeCap 18
4.1.2 MET-Meme 20
4.2 資料集分析 22
4.2.1 標籤說明與實驗設置 23
4.2.2 Sentiment Analysis實驗結果 24
4.2.3 Intention Detection實驗結果 25
4.2.4 討論 25
4.3 評估指標 26
4.4 實驗環境與設置 27
4.5 純檢索實驗結果 28
4.5.1 MemeCap實驗結果 28
4.5.2 MET-Meme實驗結果 29
4.5.3 討論與小結 30
4.6 生成+檢索實驗結果 31
4.6.1 MemeCap實驗結果 31
4.6.2 MET-Meme實驗結果 32
4.6.3 討論與小結 33
第五章 結論與建議 34
5.1 結論 34
5.2 研究限制 35
5.2.1 資料集 35
5.2.2 評估指標 35
5.3 未來展望 36
參考文獻 37
參考文獻 References
[1] Rasiwasia, N., Costa Pereira, J., Coviello, E., Doyle, G., Lanckriet, G. R., Levy, R., & Vasconcelos, N. (2010, October). A new approach to cross-modal multimedia retrieval. In Proceedings of the 18th ACM international conference on Multimedia (pp. 251-260).
[2] Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., & Ng, A. Y. (2011, January). Multimodal deep learning. In ICML.
[3] Guo, J., Fan, Y., Pang, L., Yang, L., Ai, Q., Zamani, H., ... & Cheng, X. (2020). A deep look into neural ranking models for information retrieval. Information Processing & Management, 57(6), 102067.
[4] Latif, A., Rasheed, A., Sajid, U., Ahmed, J., Ali, N., Ratyal, N. I., ... & Khalil, T. (2019). Content-based image retrieval and feature extraction: a comprehensive review. Mathematical Problems in Engineering, 2019.
[5] Dang-Nguyen, D. T., Piras, L., Giacinto, G., Boato, G., & Natale, F. G. D. (2017). Multimodal retrieval with diversification and relevance feedback for tourist attraction images. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 13(4), 1-24.
[6] Milo, T., Somech, A., & Youngmann, B. (2019, April). Simmeme: A search engine for internet memes. In 2019 IEEE 35th International Conference on Data Engineering (ICDE) (pp. 974-985). IEEE.s
[7] Singh, A., Bacchuwar, K., & Bhasin, A. (2012). A survey of OCR applications. International Journal of Machine Learning and Computing, 2(3), 314.
[8] Perez-Martin, J., Bustos, B., & Saldana, M. (2020). Semantic Search of Memes on Twitter. arXiv preprint arXiv:2002.01462.
[9] Atrey, P. K., Hossain, M. A., El Saddik, A., & Kankanhalli, M. S. (2010). Multimodal fusion for multimedia analysis: a survey. Multimedia systems, 16, 345-379.
[10] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
[11] Li, Z., Liu, F., Yang, W., Peng, S., & Zhou, J. (2021). A survey of convolutional neural networks: analysis, applications, and prospects. IEEE transactions on neural networks and learning systems, 33(12), 6999-7019.
[12] Lu, J., Batra, D., Parikh, D., & Lee, S. (2019). Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. arXiv preprint arXiv:1908.02265.
[13] Kiela, D., Firooz, H., Mohan, A., Goswami, V., Singh, A., Ringshia, P., & Testuggine, D. (2020). The hateful memes challenge: Detecting hate speech in multimodal memes. arXiv preprint arXiv:2005.04790.
[14] Muennighoff, N. (2020). Vilio: State-of-the-art Visio-Linguistic Models applied to Hateful Memes. arXiv preprint arXiv:2012.07788.
[15] Vyalla, S. R., & Udandarao, V. (2020). Memeify: A large-scale meme generation system. In Proceedings of the 7th ACM IKDD CoDS and 25th COMAD (pp. 307-311).
[16] Peirson V, A. L., & Tolunay, E. M. (2018). Dank learning: Generating memes using deep neural networks. arXiv preprint arXiv:1806.04510.
[17] Lopes, J. P., Cunha, J. M., & Martins, P. Computational Creativity in Meme Generation: A Multimodal Approach.
[18] Wang, L., Zhang, Q., Kim, Y., Wu, R., Jin, H., Deng, H., ... & Kim, C. H. (2021). Automatic Chinese meme generation using deep neural networks. IEEE Access, 9, 152657-152667.
[19] Shimomoto, E. K., Souza, L. S., Gatto, B. B., & Fukui, K. (2019, May). News2meme: An automatic content generator from news based on word subspaces from text and image. In 2019 16th International Conference on Machine Vision Applications (MVA) (pp. 1-6). IEEE.
[20] Sadasivam, A., Gunasekar, K., Davulcu, H., & Yang, Y. (2020). memebot: Towards automatic image meme generation. arXiv preprint arXiv:2004.14571.
[21] Wang, H., & Lee, R. K. W. (2024, May). MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation. In Proceedings of the ACM on Web Conference 2024 (pp. 4642-4652).
[22] Hwang, E., & Shwartz, V. (2023). Memecap: A dataset for captioning and interpreting memes. arXiv preprint arXiv:2305.13703.
[23] Xu, B., Li, T., Zheng, J., Naseriparsa, M., Zhao, Z., Lin, H., & Xia, F. (2022, July). Met-meme: A multimodal meme dataset rich in metaphors. In Proceedings of the 45th international ACM SIGIR conference on research and development in information retrieval (pp. 2887-2899).
[24] Kougia, V., Fetzel, S., Kirchmair, T., Çano, E., Baharlou, S. M., Sharifzadeh, S., & Roth, B. (2023, August). Memegraphs: Linking memes to knowledge graphs. In International Conference on Document Analysis and Recognition (pp. 534-551). Cham: Springer Nature Switzerland.
電子全文 Fulltext
本電子全文僅授權使用者為學術研究之目的,進行個人非營利性質之檢索、閱讀、列印。請遵守中華民國著作權法之相關規定,切勿任意重製、散佈、改作、轉貼、播送,以免觸法。
論文使用權限 Thesis access permission:自定論文開放時間 user define
開放時間 Available:
校內 Campus:開放下載的時間 available 2027-08-08
校外 Off-campus:開放下載的時間 available 2027-08-08

您的 IP(校外) 位址是 18.97.9.175
現在時間是 2026-08-11
論文校外開放下載的時間是 2027-08-08

Your IP address is 18.97.9.175
The current date is 2026-08-11
This thesis will be available to you on 2027-08-08.

紙本論文 Printed copies
紙本論文的公開資訊在102學年度以後相對較為完整。如果需要查詢101學年度以前的紙本論文公開資訊,請聯繫圖資處紙本論文服務櫃台。如有不便之處敬請見諒。
開放時間 available 2027-08-08

QR Code