Subreddit Engagement Velocity and Chapter-Level Reader Retention in English Webtoons: A Survival Analysis with Window-Decay and Power Diagnostics

Authors

  • Qixuan Zhang Shenzhen College of International Education, Shenzhen, China

DOI:

https://doi.org/10.6911/

Keywords:

Survival analysis · proportional hazards · reader retention · webtoon · Reddit · cross-platform fan signals · pre-registered null · power analysis.

Abstract

LINE Webtoon hosts hundreds of long-running English-language webcomics, but quantitative work in this area predicts initial popularity rather than chapter-by-chapter reader retention. We assemble a panel of 117 English LINE Webtoon Originals from the public catalog and join them to AniList records. For each title we collect per-episode public like counts and harvest 50,103 Reddit submissions and comments mentioning the title across four webtoon-related subreddits, using the PullPush mirror of Pushshift. We estimate Cox proportional-hazards models of a chapter-level dropoff event, defined as a sustained ≥50% decline in chapter likes relative to the title’s first-chapter baseline, with title-level frailty. Adding early-window Reddit engagement to a baseline of platform-internal covariates changes the concordance index by ΔC ̂=0.027 (95% cluster bootstrap CI: [-0.014,0.106]). The point estimate is directionally consistent with Reddit signals adding modest predictive value, but the cluster-bootstrap interval spans zero and we therefore cannot reject the pre-registered null at the chosen ΔC≥0.02 threshold. A post-hoc empirical-power analysis that simulates from the fitted full model (true effect ≈0.027) shows the threshold-clearance test has approximately zero power at n=56, indicating that the wide interval is structural rather than evidentiary. A window-length sweep further reveals that the incremental Reddit lift is positive only when restricted to the first one to two weeks after launch (ΔC ̂=+0.027 at W=1 week) and turns negative for W≥4 weeks, suggesting that off-platform discussion volume becomes redundant with platform engagement once aggregated over longer windows. Robustness across discrete-time hazard, sensitivity-grid, and Bayesian-prior specifications is consistent. We also report a null result for Tumblr Fandometrics: across 83 weeks of weekly anime-manga charts, none of the top-25 fandoms matched our cohort. The cohort, the chapter-likes panel, the Reddit harvest, and the analysis code are released for replication.

Downloads

Download data is not yet available.

References

[1] Baumgartner, J., Zannettou, S., Keegan, B., Squire, M., & Blackburn, J. (2020). The pushshift reddit dataset. Proceedings of the International AAAI Conference on Web and Social Media, 14, 830–839. https://doi.org/10.1609/icwsm.v14i1.7347.

[2] Davidson, B. I., & Joinson, A. N. (2025). Reddit in scholarly reception: A bibliometric assessment of the front page of the internet. Quality & Quantity. https://doi.org/10.1007/s11135-025-02416-z

[3] Cox, D. R. (1972). Regression models and life-tables. Journal of the Royal Statistical Society: Series B (Methodological), 34(2), 187–202. https://doi.org/10.1111/j.2517-6161.1972.tb00899.x

[4] Kalbfleisch, J. D., & Schaubel, D. E. (2023). Fifty years of the cox model. Annual Review of Statistics and Its Application, 10, 1–23. https://doi.org/10.1146/annurev-statistics-033021-014043

[5] Cho, H., Adkins, D., & Long, A. K. (2025). Understanding the reader demographics of an emerging online reading platform, webtoon. Journal of Documentation, 81(2), 351–368. https://doi.org/10.1108/JD-03-2024-0069

[6] Kim, J.-H., & Yu, J. (2019). Platformizing webtoons: The impact on creative and digital labor in South Korea. Social Media + Society, 5(4), 1–11. https://doi.org/10.1177/2056305119880174

[7] Yecies, B., Shim, A., Yang, J., & Zhong, P. Y. (2020). Global transcreators and the extension of the Korean webtoon IP-engine. Media, Culture & Society, 42(1), 40–57. https://doi.org/10.1177/0163443719867277

[8] Sugishita, K., & Masuda, N. (2023). Social network analysis of manga: Similarities to real-world social networks and trends over decades. Applied Network Science, 8, 79. https://doi.org/10.1007/s41109-023-00604-0

[9] Cho, H., Adkins, D., & Long, A. K. (2022). “I only wish that I had had that growing up”: Understanding webtoon’s appeals and characteristics as an emerging reading platform. Proceedings of the Association for Information Science and Technology, 59, 448–453. https://doi.org/10.1002/pra2.603

[10] Shim, A., Yecies, B., Ren, X., & Wang, D. (2020). Cultural intermediation and the basis of trust among webtoon and webnovel communities. Information, Communication & Society, 23(6), 833–848. https://doi.org/10.1080/1369118X.2020.1751865

[11] Yoon, S. (2024). Webtoons, desperately seeking viewers: Interactive creativity in social media platforms and cultural appropriation of global media production. Social Media + Society, 10(4), 1–12. https://doi.org/10.1177/20563051241292577

[12] Austin, P. C. (2017). A tutorial on multilevel survival analysis: Methods, models and applications. International Statistical Review, 85(2), 185–203. https://doi.org/10.1111/insr.12214

[13] Therneau, T. M., & Grambsch, P. M. (2000). Modeling survival data: Extending the cox model. Springer. https://doi.org/10.1007/978-1-4757-3294-8

[14] Webb, A., & Ma, J. (2023). Cox models with time-varying covariates and partly-interval censoring—A maximum penalised likelihood approach. Statistics in Medicine, 42(6), 815–833. https://doi.org/10.1002/sim.9645

[15] Pinto, H., Almeida, J. M., & Gonçalves, M. A. (2013). Using early view patterns to predict the popularity of YouTube videos. Proceedings of the Sixth ACM International Conference on Web Search and Data Mining (WSDM), 365–374. https://doi.org/10.1145/2433396.2433443

[16] Asur, S., & Huberman, B. A. (2010). Predicting the future with social media. Proceedings of the 2010 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT), 1, 492–499. https://doi.org/10.1109/WI-IAT.2010.63

[17] Cha, M., Kwak, H., Rodriguez, P., Ahn, Y.-Y., & Moon, S. (2009). Analyzing the video popularity characteristics of large-scale user generated content systems. IEEE/ACM Transactions on Networking, 17(5), 1357–1370. https://doi.org/10.1109/TNET.2008.2011358

[18] Szabo, G., & Huberman, B. A. (2010). Predicting the popularity of online content. Communications of the ACM, 53(8), 80–88. https://doi.org/10.1145/1787234.1787254

[19] Figueiredo, F., Almeida, J. M., Gonçalves, M. A., & Benevenuto, F. (2014). On the dynamics of social media popularity: A YouTube case study. ACM Transactions on Internet Technology, 14(4), 24:1–24:23. https://doi.org/10.1145/2665065

[20] Wang, D., & Yecies, B. (2023). Collaborative cultural intermediation and communitainment value in the creator economy: The expanding case of webtoons. Global Media and China. https://doi.org/10.1177/27523543231219450

[21] Tan, W. H., & Chen, F. (2021). Predicting the popularity of tweets using internal and external knowledge: An empirical bayes type approach. AStA Advances in Statistical Analysis, 105(2), 335–352. https://doi.org/10.1007/s10182-021-00390-z

[22] Gaffney, D., & Matias, J. N. (2018). Caveat emptor, computational social science: Large-scale missing data in a widely-published reddit corpus. PLOS ONE, 13(7), e0200162. https://doi.org/10.1371/journal.pone.0200162

[23] Proferes, N., Jones, N., Gilbert, S., Fiesler, C., & Zimmer, M. (2021). Studying reddit: A systematic overview of disciplines, approaches, methods, and ethics. Social Media + Society, 7(2), 1–14. https://doi.org/10.1177/20563051211019004

[24] Jasser, J., Garibay, I., Scheinert, S., & Mantzaris, A. V. (2022). Controversial information spreads faster and further than non-controversial information in reddit. Journal of Computational Social Science, 5(1), 111–122. https://doi.org/10.1007/s42001-021-00121-z

[25] Murtfeldt, R., Alterman, N., Kahveci, I., & West, J. D. (2024). RIP Twitter API: A eulogy to its vast research contributions. arXiv Preprint arXiv:2404.07340. https://doi.org/10.48550/arXiv.2404.07340

[26] Harrell, F. E., Lee, K. L., & Mark, D. B. (1996). Multivariable prognostic models: Issues in developing models, evaluating assumptions and adequacy, and measuring and reducing errors. Statistics in Medicine, 15(4), 361–387. https://doi.org/10.1002/(SICI)1097-0258(19960229)15:4<361::AID-SIM168>3.0.CO;2-4

[27] Cheng, J., Adamic, L., Dow, P. A., Kleinberg, J. M., & Leskovec, J. (2014). Can cascades be predicted? Proceedings of the 23rd International Conference on World Wide Web (WWW), 925–936. https://doi.org/10.1145/2566486.2567997

[28] Rizoiu, M.-A., Xie, L., Sanner, S., Cebrian, M., Yu, H., & Van Hentenryck, P. (2017). Expecting to be HIP: Hawkes intensity processes for social media popularity. Proceedings of the 26th International Conference on World Wide Web, 735–744. https://doi.org/10.1145/3038912.3052650

[29] Mishra, S., Rizoiu, M.-A., & Xie, L. (2016). Feature driven and point process approaches for popularity prediction. Proceedings of the 25th ACM International Conference on Information and Knowledge Management (CIKM), 1069–1078. https://doi.org/10.1145/2983323.2983812

[30] Martin, T., Hofman, J. M., Sharma, A., Anderson, A., & Watts, D. J. (2016). Exploring limits to prediction in complex social systems. Proceedings of the 25th International Conference on World Wide Web (WWW), 683–694. https://doi.org/10.1145/2872427.2883001

[31] Shulman, B., Sharma, A., & Cosley, D. (2016). Predictability of popularity: Gaps between prediction and understanding. Proceedings of the International AAAI Conference on Web and Social Media, 10(1), 348–357. https://doi.org/10.1609/icwsm.v10i1.14748

[32] Yang, J., & Leskovec, J. (2011). Patterns of temporal variation in online media. Proceedings of the Fourth ACM International Conference on Web Search and Data Mining, 177–186. https://doi.org/10.1145/1935826.1935863

[33] Ng, K. W., Mubang, F., Hall, L. O., Skvoretz, J., & Iamnitchi, A. (2023). Experimental evaluation of baselines for forecasting social media time series. EPJ Data Science, 12(1), 8. https://doi.org/10.1140/epjds/s13688-023-00383-9

[34] Wiegrebe, S., Kopper, P., Sonabend, R., Bischl, B., & Bender, A. (2024). Deep learning for survival analysis: A review. Artificial Intelligence Review, 57(3), 65. https://doi.org/10.1007/s10462-023-10681-3

[35] Long, S., Lucey, B. M., Xie, Y., & Yarovaya, L. (2023). “I just like the stock”: The role of reddit sentiment in the GameStop share rally. Financial Review, 58(1), 19–37. https://doi.org/10.1111/fire.12328

[36] Betzer, A., & Harries, J. P. (2022). How online discussion board activity affects stock trading: The case of GameStop. Financial Markets and Portfolio Management, 36, 443–472. https://doi.org/10.1007/s11408-022-00407-w

[37] Nadiri, A., & Takes, F. W. (2022). A large-scale temporal analysis of user lifespan durability on the reddit social media platform. Companion Proceedings of the Web Conference 2022 (WWW ’22), 677–685. https://doi.org/10.1145/3487553.3524699

[38] Savela, N., Pellert, M., Latikka, R., Bergdahl, J., Garcia, D., & Oksanen, A. (2025). Affective, cognitive, and contextual cues in reddit posts on artificial intelligence. Journal of Computational Social Science, 8(1), 6. https://doi.org/10.1007/s42001-024-00335-x

[39] Halawa, S., Greene, D., & Mitchell, J. (2014). Dropout prediction in MOOCs using learner activity features. Proceedings of the European MOOC Stakeholder Summit, 58–65.

[40] Ameri, S., Fard, M. J., Chinnam, R. B., & Reddy, C. K. (2016). Survival analysis based framework for early prediction of student dropouts. Proceedings of the 25th ACM International Conference on Information and Knowledge Management (CIKM), 903–912. https://doi.org/10.1145/2983323.2983351

[41] Chi, Z., Zhang, S., & Shi, L. (2023). Analysis and prediction of MOOC learners’ dropout behavior. Applied Sciences, 13(2), 1068. https://doi.org/10.3390/app13021068

[42] Borrella, I., Caballero-Caballero, S., & Ponce-Cueto, E. (2022). Taking action to reduce dropout in MOOCs: Tested interventions. Computers & Education, 179, 104412. https://doi.org/10.1016/j.compedu.2021.104412

[43] Prenkaj, B., Velardi, P., Stilo, G., Distante, D., & Faralli, S. (2020). A survey of machine learning approaches for student dropout prediction in online courses. ACM Computing Surveys, 53(3), 57:1–57:34. https://doi.org/10.1145/3388792

[44] Reddy, S., Lazarova, M., Yu, Y., & Jones, R. (2021). Modeling language usage and listener engagement in podcasts. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, 632–643. https://doi.org/10.18653/v1/2021.acl-long.52

[45] Pan, F., Huang, B., Zhang, C., et al. (2022). A survival analysis based volatility and sparsity modeling network for student dropout prediction. PLOS ONE, 17(5), e0267138. https://doi.org/10.1371/journal.pone.0267138

[46] Kaplan, E. L., & Meier, P. (1958). Nonparametric estimation from incomplete observations. Journal of the American Statistical Association, 53(282), 457–481. https://doi.org/10.1080/01621459.1958.10501452

[47] Begun, A., Kulinskaya, E., & Ncube, M. (2023). A double-cox model for non-proportional hazards survival analysis with frailty. Statistics in Medicine, 42(18), 3114–3127. https://doi.org/10.1002/sim.9760

[48] Ishwaran, H., Kogalur, U. B., Blackstone, E. H., & Lauer, M. S. (2008). Random survival forests. The Annals of Applied Statistics, 2(3), 841–860. https://doi.org/10.1214/08-AOAS169

[49] Davidson-Pilon, C. (2019). lifelines: Survival analysis in Python. Journal of Open Source Software, 4(40), 1317. https://doi.org/10.21105/joss.01317.

Downloads

Published

2026-09-20

Issue

Section

Articles

How to Cite

Zhang, Q. (2026). Subreddit Engagement Velocity and Chapter-Level Reader Retention in English Webtoons: A Survival Analysis with Window-Decay and Power Diagnostics. World Scientific Research Journal, 12(10), 188-209. https://doi.org/10.6911/