A Reproducible Resource-Aware Cascade for Wildlife Image Analysis Under Limited Data
DOI:
https://doi.org/10.6911/Keywords:
Camera traps; wildlife image classification; conditional inference; model cascade; confidence calibration; group-stratified cross-validation; resource-aware inference.Abstract
Selecting when to invoke a costly classifier is harder than evaluating the cost-performance curve after the fact. We examine this problem in a reproducible wildlife image-analysis prototype that uses a lightweight model for routine screening and a stronger model for selected cases. The evaluation uses 1,000 verified ENA24 images from four balanced classes. Numeric-ID burst proxies and identical 64-bit difference hashes produce 160 conservative groups. Five outer group-stratified folds keep fitting, model selection, calibration, routing validation, and testing separate. Frozen ImageNet-pretrained MobileNetV3-Small and ResNet-18 extract features for multinomial logistic-regression heads. Temperature scaling is fitted on calibration data. Each fold then selects the lowest-cost threshold whose routing-validation Macro-F1 is within 0.02 of stage 2. The pooled out-of-fold Macro-F1 is 0.89535 for stage 1 and 0.92397 for stage 2. The validation-selected cascade calls stage 2 for 3.9% of test images and reaches 0.89720, a 0.02677 gap from stage 2 that exceeds the specified tolerance. Higher fixed budgets approach stage-2 performance, but they do not change the failure of the prospective selection rule. The result points to threshold-transfer instability in this small-data setting. Acoustic or other contextual signals are a possible extension, not an experiment reported here. The resulting artifact is a reproducible visual prototype with an explicit evaluation protocol and a clearly reported negative result.
Downloads
References
[1] Swanson, A., Kosmala, M., Lintott, C., Simpson, R., Smith, A., & Packer, C. (2015). Snapshot Serengeti, high-frequency annotated camera trap images of 40 mammalian species in an African savanna. Scientific Data, 2, 150026. https://doi.org/10.1038/sdata.2015.26.
[2] Labeled Information Library of Alexandria: Biology and Conservation. (2019). ENA24-detection. LILA BC dataset page. https://lila.science/datasets/ena24detection
[3] Norouzzadeh, M. S., et al. (2018). Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning. Proceedings of the National Academy of Sciences, 115(25), E5716–E5725. https://doi.org/10.1073/pnas.1719367115
[4] Yousif, H., Yuan, J., Kays, R., & He, Z. (2017). Fast human-animal detection from highly cluttered camera-trap images using joint background modeling and deep learning classification. In 2017 IEEE International Symposium on Circuits and Systems (ISCAS) (pp. 1–4). https://doi.org/10.1109/ISCAS.2017.8050762
[5] Yousif, H., Yuan, J., Kays, R., & He, Z. (2019). Animal scanner: Software for classifying humans, animals, and empty frames in camera trap images. Ecology and Evolution, 9(4), 1578–1589. https://doi.org/10.1002/ece3.4747
[6] Teerapittayanon, S., McDanel, B., & Kung, H. T. (2016). BranchyNet: Fast inference via early exiting from deep neural networks. In 2016 23rd International Conference on Pattern Recognition (ICPR) (pp. 2464–2469). https://doi.org/10.1109/ICPR.2016.7900006
[7] Bolukbasi, T., Wang, J., Dekel, O., & Saligrama, V. (2017). Adaptive neural networks for efficient inference. In Proceedings of the 34th International Conference on Machine Learning (Vol. 70, pp. 527–536). PMLR.
[8] Wang, X., Yu, F., Dou, Z.-Y., Darrell, T., & Gonzalez, J. E. (2018). SkipNet: Learning dynamic routing in convolutional networks. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 420–436). https://doi.org/10.1007/978-3-030-01261-8_25
[9] Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (Vol. 70, pp. 1321–1330). PMLR.
[10] Nixon, J., Dusenberry, M. W., Zhang, L., Jerfel, G., & Tran, D. (2019). Measuring calibration in deep learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (pp. 38–41).
[11] Beery, S., Van Horn, G., & Perona, P. (2018). Recognition in terra incognita. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 472–489). https://doi.org/10.1007/978-3-030-01270-0_28
[12] Tabak, M. A., et al. (2020). Improving the accessibility and transferability of machine learning algorithms for identification of animals in camera trap images: MLWIC2. Ecology and Evolution, 10(19), 10374–10383. https://doi.org/10.1002/ece3.6692
[13] Howard, A., et al. (2019). Searching for MobileNetV3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 1314–1324). https://doi.org/10.1109/ICCV.2019.00140
[14] He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770–778). https://doi.org/10.1109/CVPR.2016.90
[15] Kline, J., Potlapally, A., Pillai, B., Wani, T., Patil, V., Covey, P., Katole, R., Stevens, S., Subramoni, H., Berger-Wolf, T., & Stewart, C. (2025). SmartWilds: Multimodal wildlife monitoring dataset. arXiv preprint arXiv:2509.18894 (Version 2). https://arxiv.org/abs/2509.18894v2
[16] Pillai, B., Viswapriyan, V., Stewart, C., Berger-Wolf, T., & Kline, J. (2026). Cross-modal corroboration for annotation-free wildlife monitoring. arXiv preprint arXiv:2606.21613 (Version 1). https://arxiv.org/abs/2606.21613v1
[17] Yousif, H., Kays, R., & He, Z. (2019). Dynamic programming selection of object proposals for sequence-level animal species classification in the wild. IEEE Transactions on Circuits and Systems for Video Technology.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 World Scientific Research Journal

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.




