DriftFair: Anytime-Valid Temporal Fairness Monitoring and Shift Attribution for AI-Assisted Talent Promotion Decisions

Authors

  • Zimeng Wang Brandeis University, Waltham, MA 02453, United States
  • Zifan Chen University of Pennsylvania, Philadelphia, PA 19104, United States
  • Zijian Shen Carnegie Mellon University, Pittsburgh, PA 15213, United States

DOI:

https://doi.org/10.6911/

Keywords:

Algorithmic fairness, concept drift, sequential change detection, anytime-valid inference, e-values, model monitoring, talent analytics.

Abstract

Organisations increasingly deploy machine-learned scorers to shortlist employees for promotion. Such systems are typically audited for fairness once, at deployment, and then left unchecked, even though the population they judge keeps moving. We show that this practice is unsafe in a specific and measurable way: the drift that damages fairness and the drift that damages accuracy are largely disjoint, so accuracy-oriented monitors never fire, while input-oriented monitors fire on shifts that carry no fairness consequence at all. We present DriftFair, an operational monitor with four components. (i) A change detector for group-fairness functionals built as a weighted mixture of restarted betting e-processes; because the mixture is a non-negative supermartingale started at one, Ville's inequality gives a time-uniform false-alarm guarantee over an unbounded monitoring horizon, together with an explicit tolerance band that also absorbs the estimation error of the certified reference. (ii) An additive decomposition of any observed fairness drift into within-group covariate shift, latent (concept) shift and group-composition shift, obtained by per-group density-ratio reweighting, plus an exact Shapley attribution that names the responsible feature. (iii) A multi-cohort dashboard with online false-discovery-rate control via e-values. (iv) An attribution-driven maintenance policy. On a real promotion data set of 54,808 employee records, DriftFair is the only monitor among seven that combines a zero false-alarm rate on fairness-irrelevant drift with reliable detection of harmful drift; its shift attribution reaches 0.85 macro accuracy against 0.40 for a Kolmogorov-Smirnov input monitor, and its Shapley term recovers the perturbed feature with an attributed magnitude of 0.185 versus at most 0.021 for every other feature. Monitor-triggered per-group threshold recalibration cuts excess-disparity exposure by 48% relative to no maintenance while consuming 1,100 fresh labels, about one forty-ninth of the cheapest periodic schedule, and matches an oracle that knows the true change points. We further report a real chronological stream and a quantified data-sufficiency limit for small cohorts.

Downloads

Download data is not yet available.

References

[1] Hardt, M., Price, E., & Srebro, N. (2016). Equality of opportunity in supervised learning. In Advances in Neural Information Processing Systems (pp. 3315–3323).

[2] Dwork, C., Hardt, M., Pitassi, T., Reingold, O., & Zemel, R. (2012). Fairness through awareness. In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference (pp. 214–226).

[3] Chouldechova, A. (2017). Fair prediction with disparate impact: A study of bias in recidivism prediction instruments. Big Data, 5(2), 153–163.

[4] Kleinberg, J., Mullainathan, S., & Raghavan, M. (2017). Inherent trade-offs in the fair determination of risk scores. In Proceedings of the 8th Innovations in Theoretical Computer Science Conference (pp. 43:1–43:23).

[5] Kamiran, F., & Calders, T. (2012). Data preprocessing techniques for classification without discrimination. Knowledge and Information Systems, 33(1), 1–33.

[6] Zafar, M. B., Valera, I., Gomez-Rodriguez, M., & Gummadi, K. P. (2017). Fairness constraints: Mechanisms for fair classification. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (pp. 962–970).

[7] Agarwal, A., Beygelzimer, A., Dudik, M., Langford, J., & Wallach, H. (2018). A reductions approach to fair classification. In Proceedings of the 35th International Conference on Machine Learning (pp. 60–69).

[8] Pleiss, G., Raghavan, M., Wu, F., Kleinberg, J., & Weinberger, K. Q. (2017). On fairness and calibration. In Advances in Neural Information Processing Systems (pp. 5680–5689).

[9] Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 115:1–115:35.

[10] Bellamy, R. K. E., Dey, K., Hind, M., Hoffman, S. C., Houde, S., Kalra, K., Karmahapatra, P., Martino, R., Mehta, S., Mojsilovic, A., Nagar, S., Ramamurthy, K. N., Richards, J., Saha, D., Sattigeri, P., Smith, C., Varshney, K. R., & Zhang, Y. (2019). AI Fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias. IBM Journal of Research and Development, 63(4/5), 4:1–4:15.

[11] Bird, S., Dudik, M., Edgar, R., Horn, B., Lutz, R., Milan, V., Sameki, M., Wallach, H., & Walker, K. (2020). Fairlearn: A toolkit for assessing and improving fairness in AI. Microsoft Research, Technical Report MSR-TR-2020-32.

[12] Saleiro, P., Kuester, B., Hinkson, L., London, J., Stevens, A., Anisfeld, A., Rodolfa, K. T., & Ghani, R. (2018). Aequitas: A bias and fairness audit toolkit. https://arxiv.org/abs/1811.05577

[13] Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Bass, L., Bateman, C., Brinton, H., Hua, R., Hurst, K., Kim, L., Lee, J., & Viegas, F. (2019). Model cards for model reporting. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (pp. 220–229).

[14] Raghavan, M., Barocas, S., Kleinberg, J., & Levy, K. (2020). Mitigating bias in algorithmic hiring: Evaluating claims and practices. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (pp. 469–481).

[15] Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and auditing fair algorithms: A case study in candidate screening. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (pp. 666–677).

[16] Barocas, S., & Selbst, A. D. (2016). Big data’s disparate impact. California Law Review, 104(3), 671–732.

[17] Angwin, J., Larson, J., Mattu, S., & Kirchner, L. (2016, May 23). Machine bias. ProPublica.

[18] Gama, J., Zliobaite, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), 44:1–44:37.

[19] Lu, J., Liu, A., Dong, F., Gu, F., Gama, J., & Zhang, G. (2019). Learning under concept drift: A review. IEEE Transactions on Knowledge and Data Engineering, 31(12), 2346–2363.

[20] Bifet, A., & Gavalda, R. (2007). Learning from time-changing data with adaptive windowing. In Proceedings of the SIAM International Conference on Data Mining (pp. 443–448).

[21] Gama, J., Medas, P., Castillo, G., & Rodrigues, P. (2004). Learning with drift detection. In Advances in Artificial Intelligence (pp. 286–295).

[22] Page, E. S. (1954). Continuous inspection schemes. Biometrika, 41(1/2), 100–115.

[23] Rabanser, S., Guennemann, S., & Lipton, Z. C. (2019). Failing loudly: An empirical study of methods for detecting dataset shift. In Advances in Neural Information Processing Systems (pp. 1396–1408).

[24] Gretton, A., Borgwardt, K. M., Rasch, M. J., Schoelkopf, B., & Smola, A. (2012). A kernel two-sample test. Journal of Machine Learning Research, 13, 723–773.

[25] Lipton, Z. C., Wang, Y.-X., & Smola, A. (2018). Detecting and correcting for label shift with black box predictors. In Proceedings of the 35th International Conference on Machine Learning (pp. 3122–3130).

[26] Ding, F., Hardt, M., Miller, J., & Schmidt, L. (2021). Retiring Adult: New datasets for fair machine learning. In Advances in Neural Information Processing Systems (pp. 6478–6490).

[27] Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Kreuger, J., & Dennison, D. (2015). Hidden technical debt in machine learning systems. In Advances in Neural Information Processing Systems (pp. 2503–2511).

[28] Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. In Proceedings of the IEEE International Conference on Big Data (pp. 1123–1132).

[29] Rhee, M., Zou, J., Mo, T., Teng, D., & Yang, J.-S. (2026). Enhancing web search agents with self-play contrastive fine-tuning. IEEE Access. https://doi.org/10.1109/ACCESS.2026.3717594

[30] Fan, H., Jiao, Y., Wang, M., Chen, J., & Yang, S. (2026). Fraud learns too: Continual graph learning under strategic adversarial drift in dynamic networks. Scientific Reports. https://doi.org/10.1038/s41598-026-60997-7

[31] D’Amour, A., Srinivasan, H., Atwood, J., Baljekar, P., Sculley, D., & Halpern, Y. (2020). Fairness is not static: Deeper understanding of long term fairness via simulation studies. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (pp. 525–534).

[32] Liu, L. T., Dean, S., Rolf, E., Simchowitz, M., & Hardt, M. (2018). Delayed impact of fair machine learning. In Proceedings of the 35th International Conference on Machine Learning (pp. 3150–3158).

[33] Zhang, W., & Ntoutsi, E. (2019). FAHT: An adaptive fairness-aware decision tree classifier. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (pp. 1480–1486).

[34] Iosifidis, V., & Ntoutsi, E. (2020). FABBOO: Online fairness-aware learning under class imbalance. In Discovery Science (pp. 159–174).

[35] Zhang, W., Bifet, A., Zhang, X., Weiss, J. C., & Nejdl, W. (2021). FARF: A fair and adaptive random forests classifier. In Proceedings of the Pacific-Asia Conference on Knowledge Discovery and Data Mining (pp. 245–256).

[36] Ghosh, A., Shanbhag, A., & Wilson, C. (2022). FairCanary: Rapid continuous explainable fairness. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (pp. 307–316).

[37] Waudby-Smith, I., & Ramdas, A. (2024). Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society Series B, 86(1), 1–27.

[38] Ramdas, A., Grunwald, P., Vovk, V., & Shafer, G. (2023). Game-theoretic statistics and safe anytime-valid inference. Statistical Science, 38(4), 576–601.

[39] Vovk, V., & Wang, R. (2021). E-values: Calibration, combination and applications. The Annals of Statistics, 49(3), 1736–1754.

[40] Wang, R., & Ramdas, A. (2022). False discovery rate control with e-values. Journal of the Royal Statistical Society Series B, 84(3), 822–852.

[41] Xu, Z., & Ramdas, A. (2024). Online multiple testing with e-values. In Proceedings of the 27th International Conference on Artificial Intelligence and Statistics (pp. 3997–4005).

[42] Shin, J., Ramdas, A., & Rinaldo, A. (2023). E-detectors: A nonparametric framework for sequential change detection. The New England Journal of Statistics in Data Science, 2(2), 229–260.

[43] Chugg, B., Cortes-Gomez, S., Wilder, B., & Ramdas, A. (2023). Auditing fairness by betting. In Advances in Neural Information Processing Systems.

[44] Cherian, J. J., & Candes, E. J. (2024). Statistical inference for fairness auditing. Journal of Machine Learning Research, 25, 1–49.

[45] Maneriker, P., Burley, C., & Parthasarathy, S. (2023). Online fairness auditing through iterative refinement. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 1665–1676).

[46] Chen, Z., Wang, M., Zeng, Z., & Ping, W. (2026). Uncertainty-aware financial forecasting: Leveraging conformal prediction for risk-adjusted models. IEEE Access, 14, 90053–90077. https://doi.org/10.1109/ACCESS.2026.3702666

[47] Liang, Y., Jiao, Y., Ping, W., Fan, H., & Han, X. (2026). Adaptive event-driven labeling: A neuro-symbolic multiagent framework for causal inference in non-stationary time series. IEEE Access, 14, 104468–104482. https://doi.org/10.1109/ACCESS.2026.3709267

[48] European Parliament and Council of the European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union.

Downloads

Published

2026-09-17

Issue

Section

Articles

How to Cite

Wang, Z., Chen, Z., & Shen, Z. (2026). DriftFair: Anytime-Valid Temporal Fairness Monitoring and Shift Attribution for AI-Assisted Talent Promotion Decisions. World Scientific Research Journal, 12(10), 32-47. https://doi.org/10.6911/