A Hybrid Machine-Learning Framework for Environmental Impact Prediction and Candidate Option Ranking

Authors

  • Yiqing Liao The Experimental High School Attached to Beijing Normal University, Beijing 100032, China
  • Zimeng Jiang The Experimental High School Attached to Beijing Normal University, Beijing 100032, China
  • Ruohan Ma The Experimental High School Attached to Beijing Normal University, Beijing 100032, China
  • Xinran Li The Experimental High School Attached to Beijing Normal University, Beijing 100032, China

DOI:

https://doi.org/10.6911/

Keywords:

Environmental impact prediction; LightGBM; AHM-PCA-FCE; maximal information coefficient; candidate option ranking; intelligent decision support.

Abstract

Large-scale activities create environmental loads across many dimensions over short periods, including venue power demand, visitor mobility, accommodation resource use, and post-activity recovery streams. Candidate option screening is treated here as a problem of computational prediction and multi-criteria ranking. A life-cycle assessment encoder organizes the target activity into five operational phases and thirteen pressure-oriented features. A six-variable context layer describes population scale, transport-network length, renewable-energy share, climate signal, service capacity, and digital service-readiness proxy. The Analytic Hierarchy Model and Principal Component Analysis are coupled to obtain hybrid feature weights, Fuzzy Comprehensive Evaluation converts weighted indicators into Environmental Impact Factor scores, LightGBM predicts scores for future or unobserved candidate options, and the Maximal Information Coefficient detects nonlinear relationships between the score and contextual features. A verification module tests leave-one-sample-out prediction error, rank stability, residual distribution, and feature-ablation performance. The three dominant computational features are February air-arrival volume, accommodation energy use per guest-night, and expected electrical load peak. Among historical cases, Indianapolis, Pontiac, and Palo Alto show the lowest Environmental Impact Factor values, while Las Vegas and Miami-related cases form the high-pressure group. For the 2029 candidate set, Seattle obtains the lowest predicted score, followed by Nashville and Kansas City. The multi-scenario extension confirms the applicability of this algorithmic pipeline to larger evaluation settings after feature re-encoding.

Downloads

Download data is not yet available.

References

[1] ISO. (2006). Environmental management—Life cycle assessment—Principles and framework (ISO 14040). International Organization for Standardization.

[2] ISO. (2006). Environmental management—Life cycle assessment—Requirements and guidelines (ISO 14044). International Organization for Standardization.

[3] Saaty, T. L. (1980). The analytic hierarchy process. McGraw-Hill.

[4] Jolliffe, I. T., & Cadima, J. (2016). Principal component analysis: A review and recent developments. Philosophical Transactions of the Royal Society A, 374(2065), 20150202. https://doi.org/10.1098/rsta.2015.0202

[5] Zadeh, L. A. (1965). Fuzzy sets. Information and Control, 8(3), 338–353. https://doi.org/10.1016/S0019-9958(65)90241-X

[6] Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 30, 3146–3154.

[7] Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189–1232. https://doi.org/10.1214/aos/1013203451

[8] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. https://doi.org/10.1145/2939672.2939785

[9] Reshef, D. N., Reshef, Y. A., Finucane, H. K., Grossman, S. R., McVean, G., Turnbaugh, P. J., Lander, E. S., Mitzenmacher, M., & Sabeti, P. C. (2011). Detecting novel associations in large data sets. Science, 334(6062), 1518–1524. https://doi.org/10.1126/science.1205438

[10] Chicco, D., Warrens, M. J., & Jurman, G. (2021). The coefficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analysis evaluation. PeerJ Computer Science, 7, e623. https://doi.org/10.7717/peerj-cs.623

[11] Greenacre, M., Groenen, P. J. F., Hastie, T., d'Enza, A. I., Markos, A., & Tuzhilina, E. (2022). Principal component analysis. Nature Reviews Methods Primers, 2(1), 100. https://doi.org/10.1038/s43586-022-00090-6

[12] Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30, 4765–4774.

Downloads

Published

2026-09-21

Issue

Section

Articles

How to Cite

Liao, Y., Jiang, Z., Ma, R., & Li, X. (2026). A Hybrid Machine-Learning Framework for Environmental Impact Prediction and Candidate Option Ranking. World Scientific Research Journal, 12(10), 223-231. https://doi.org/10.6911/