論文使用權限 Thesis access permission:自定論文開放時間 user define
開放時間 Available:
校內 Campus:開放下載的時間 available 2029-07-15
校外 Off-campus:開放下載的時間 available 2029-07-15
論文名稱 Title |
多模態分析直播觀眾的追隨行為 A Multimodal Analysis of Streaming Viewer Follow Behavior |
||
系所名稱 Department |
|||
畢業學年期 Year, semester |
語文別 Language |
||
學位類別 Degree |
頁數 Number of pages |
61 |
|
研究生 Author |
|||
指導教授 Advisor |
|||
召集委員 Convenor |
|||
口試委員 Advisory Committee |
|||
口試日期 Date of Exam |
2024-06-12 |
繳交日期 Date of Submission |
2024-07-15 |
關鍵字 Keywords |
多模態深度學習、傳播理論、直播預測、模態融合、注意力機制 Multimodal Deep Learning, Communication Model, Live Streaming, Modal Fusion, Cross-Attention |
||
統計 Statistics |
本論文已被瀏覽 490 次,被下載 0 次 The thesis/dissertation has been browsed 490 times, has been downloaded 0 times. |
中文摘要 |
直播已成為一種廣受歡迎的社交媒體體驗,允許內容創作者與觀眾進行即時互動,這種新興的線上服務已在主要社交媒體平台上廣受歡迎。其中,Twitch在2014年起就已經成為全球最大的遊戲直播平台,每月有 1700 萬名直播主在 Twitch 上進行即時直播,顯示 Twitch 的人氣和持續增長。在觀看直播主內容的同時,觀眾積極參與創作過程。他們與直播主進行即時互動,提供回饋並影響直播的內容。直播主和觀眾之間的這種動態互動是直播體驗的一個獨特的特徵。我們選取Valorant這款第一人稱射擊遊戲中最多觀看次數的影片,資料範圍涵蓋一個月,共計6550分鐘的影片進行分析。 本研究採用 Shannon and Weaver 的傳播理論作為框架,並從相關研究中方法定義出Twitch直播中三種傳播角色:直播主、內容和觀眾的特徵。本研究的資料來源於Twitch API、TwitchDownloader 及網路爬蟲,並將窗格設定為10分鐘,以此對齊模態。我們提出了一個模型,採用混合融合文字、聲音和視覺四種模態,並加入交叉注意力機制。該模型的MAE 為 11.546 、RMSE 為 18.225。提供直播主在直播時,控制臉部情緒表現、增加觀眾以及控制留言的社區感的建議。 |
Abstract |
Live streaming has become very popular on social media, allowing creators to interact with audiences in real-time. Twitch is the biggest game live-streaming platform, with 17 million broadcasters streaming live every month, showing its popularity. We analyzed the most-watched videos from the popular first-person shooter game "Valorant" over one month, totaling 6,550 minutes of video. Using communication theory as a framework, we defined the characteristics of streamers, content, and audiences in live streams based on previous research. Data sources included the Twitch API, TwitchDownloader, and web crawlers with 10-minute window sizes. We propose a multimodal model that employs hybrid fusion and cross-attention to combine numeric, textual, acoustic, and visual data. This model achieved an MAE of 11.546 and RMSE of 18.225. We provide recommendations for broadcasters to control facial expressions, increase viewer engagement, and manage the sense of community in the chat during streams. |
目次 Table of Contents |
論文審定書 i 中文摘要 ii Abstract iii Chapter 1 Introduction 1 1.1 Research background 1 1.2 Research purpose 2 Chapter 2 Literature Reviews 6 2.1 Live streaming on Twitch 6 2.2 Subscription/Donation/Gifting Intention 7 2.3 The communication model 9 2.4 Information Source (Streamer) Characteristics 11 2.5 Message (Streaming Content) Characteristics 12 2.6 Destination (Viewer) Characteristics 12 2.7 Support Vector Regression 13 2.8 Multimodal fusion architectures 14 Chapter 3 Method 16 3.1 Framework 16 3.2 Data preprocessing and modal encoding 17 3.3 Fusion 27 3.4 Self-attention and cross-attention 28 3.5 Model evaluation 30 Chapter 4 Empirical Analysis 32 4.1 Data description 32 4.2 Experimental settings 34 4.3 Model comparison 35 4.4 Conclusion 40 Chapter 5 Conclusion 42 5.1 Implication for research 42 5.2 Implication for practice 43 5.3 Limitation and feature work 43 Reference 45 Appendix 50 |
參考文獻 References |
Atrey, P. K., Hossain, M. A., El Saddik, A., & Kankanhalli, M. S. (2010). Multimodal fusion for multimedia analysis: a survey. Multimedia systems, 16, 345-379. Baltrušaitis, T., Ahuja, C., & Morency, L.-P. (2018). Multimodal machine learning: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence, 41(2), 423-443. Barbieri, F., Camacho-Collados, J., Neves, L., & Espinosa-Anke, L. (2020). Tweeteval: Unified benchmark and comparative evaluation for tweet classification. arXiv preprint arXiv:2010.12421. Belova, A., He, W., & Zhong, Z. (2019). E-Sports Talent Scouting Based on Multimodal Twitch Stream Data. arXiv preprint arXiv:1907.01615. Bergstra, J., & Bengio, Y. (2012). Random search for hyper-parameter optimization. Journal of machine learning research, 13(2). Cabeza-Ramírez, L. J., Fuentes-García, F. J., & Muñoz-Fernandez, G. A. (2021). Exploring the emerging domain of research on video game live streaming in web of science: State of the art, changes and trends. International Journal of Environmental Research and Public Health, 18(6), 2917. Cawley, G. C., & Talbot, N. L. (2010). On over-fitting in model selection and subsequent selection bias in performance evaluation. The Journal of Machine Learning Research, 11, 2079-2107. Cortes, C., Mohri, M., & Rostamizadeh, A. (2012). L2 regularization for learning kernels. arXiv preprint arXiv:1205.2653. Davis, S., & Mermelstein, P. (1980). Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. IEEE transactions on acoustics, speech, and signal processing, 28(4), 357-366. Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. 2009 IEEE conference on computer vision and pattern recognition, Gros, D., Hackenholt, A., Zawadzki, P., & Wanner, B. (2018). Interactions of Twitch users and their usage behavior. Social Computing and Social Media. Technologies and Analytics: 10th International Conference, SCSM 2018, Held as Part of HCI International 2018, Las Vegas, NV, USA, July 15-20, 2018, Proceedings, Part II 10, Guarriello, N.-B. (2019). Never give up, never surrender: Game live streaming, neoliberal work, and personalized media economies. New Media & Society, 21(8), 1750-1769. Hilvert-Bruce, Z., Neill, J. T., Sjöblom, M., & Hamari, J. (2018). Social motivations of live-streaming viewer engagement on Twitch. Computers in human behavior, 84, 58-67. Huber, P. J. (1992). Robust estimation of a location parameter. In Breakthroughs in statistics: Methodology and distribution (pp. 492-518). Springer. Johnson, M. R., & Woodcock, J. (2019). ‘It’s like the gold rush’: the lives and careers of professional video game streamers on Twitch. tv. Information, Communication & Society, 22(3), 336-351. Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980. Leung, F. F., Gu, F. F., Li, Y., Zhang, J. Z., & Palmatier, R. W. (2022). Influencer marketing effectiveness. Journal of marketing, 86(6), 93-115. Li, R., Lu, Y., Ma, J., & Wang, W. (2021). Examining gifting behavior on live streaming platforms: An identity-based motivation model. Information & Management, 58(6), 103406. Li, Y., & Peng, Y. (2021). What drives gift-giving intention in live streaming? The perspectives of emotional attachment and flow experience. International Journal of Human–Computer Interaction, 37(14), 1317-1329. Lin, Y., Yao, D., & Chen, X. (2021). Happiness begets money: Emotion and engagement in live streaming. Journal of Marketing Research, 58(3), 417-438. Liu, G. H., Sun, M., & Lee, N. C.-A. (2021). How can live streamers enhance viewer engagement in eCommerce streaming? Lu, Z., Xia, H., Heo, S., & Wigdor, D. (2018). You watch, you give, and you engage: a study of live streaming practices in China. Proceedings of the 2018 CHI conference on human factors in computing systems, McFee, B., Raffel, C., Liang, D., Ellis, D. P., McVicar, M., Battenberg, E., & Nieto, O. (2015). librosa: Audio and music signal analysis in python. SciPy, Mohammad, S., Bravo-Marquez, F., Salameh, M., & Kiritchenko, S. (2018). Semeval-2018 task 1: Affect in tweets. Proceedings of the 12th international workshop on semantic evaluation, Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., & Ng, A. Y. (2011). Multimodal deep learning. Proceedings of the 28th international conference on machine learning (ICML-11), Potamianos, G., Neti, C., Luettin, J., & Matthews, I. (2004). Audio-visual automatic speech recognition: An overview. Issues in visual and audio-visual speech processing, 22, 23. Shannon, C. E. (1948). A mathematical theory of communication. The Bell system technical journal, 27(3), 379-423. Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556. Sjöblom, M., & Hamari, J. (2017). Why do people watch others play video games? An empirical study on the motivations of Twitch users. Computers in human behavior, 75, 985-996. Smola, A. J., & Schölkopf, B. (2004). A tutorial on support vector regression. Statistics and computing, 14, 199-222. Tsai, Y.-H. H., Bai, S., Liang, P. P., Kolter, J. Z., Morency, L.-P., & Salakhutdinov, R. (2019). Multimodal transformer for unaligned multimodal language sequences. Proceedings of the conference. Association for computational linguistics. Meeting, Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30. Wan, J., Lu, Y., Wang, B., & Zhao, L. (2017). How attachment influences users’ willingness to donate to content creators in social media: A socio-technical systems perspective. Information & Management, 54(7), 837-850. Xi, D., Tang, L., Chen, R., & Xu, W. (2023). A multimodal time-series method for gifting prediction in live streaming platforms. Information Processing & Management, 60(3), 103254. Zadeh, A. B., Liang, P. P., Poria, S., Cambria, E., & Morency, L.-P. (2018). Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Zhang, Y., Hua, L., Jiao, Y., Zhang, J., & Saini, R. (2023). More than watching: An empirical and experimental examination on the impacts of live streaming user-generated video consumption. Information & Management, 60(3), 103771. Zhou, J., Zhou, J., Ding, Y., & Wang, H. (2019). The magic of danmaku: A social interaction perspective of gift sending on live streaming platforms. Electronic Commerce Research and Applications, 34, 100815. |
電子全文 Fulltext |
本電子全文僅授權使用者為學術研究之目的,進行個人非營利性質之檢索、閱讀、列印。請遵守中華民國著作權法之相關規定,切勿任意重製、散佈、改作、轉貼、播送,以免觸法。 論文使用權限 Thesis access permission:自定論文開放時間 user define 開放時間 Available: 校內 Campus:開放下載的時間 available 2029-07-15 校外 Off-campus:開放下載的時間 available 2029-07-15 您的 IP(校外) 位址是 18.97.9.168 現在時間是 2026-09-15 論文校外開放下載的時間是 2029-07-15 Your IP address is 18.97.9.168 The current date is 2026-09-15 This thesis will be available to you on 2029-07-15. |
紙本論文 Printed copies |
紙本論文的公開資訊在102學年度以後相對較為完整。如果需要查詢101學年度以前的紙本論文公開資訊,請聯繫圖資處紙本論文服務櫃台。如有不便之處敬請見諒。 開放時間 available 2029-07-15 |
QR Code |