Latency Bias in Face Recognition Systems: Measurement, Statistical Evidence, and Mitigation Challenges for Eyeglass Wearers
Журнал: International Journal of Education and Management Engineering @ijeme
Статья в выпуске: 5 vol.16, 2026 года.
Бесплатный доступ
The face recognition systems (FRS) can have latency bias, with some of their inputs or subsets having disproportionate inference delays, which result in unfair quality-of-service despite apparently similar predictive performance. Although much research on fairness in face recognition research has been done on accuracy-based disparities, there has been little work on runtime behavior. This study seeks to understand whether there are differences in performance in real-time face recognition systems based on visually complex sub-groups, and whether these differences can be considered as operational fairness concerns. In this work, we investigate the latency problem of real-time FRS, with especially great attention to tail inference as the first-class fairness measure. We utilize a label-free auditing framework for conducting runtime fairness assessment, and propose the tail inference latency metric for evaluating the fairness level of the deployment. Instead of focusing on prediction scores as in prior work, our work focuses on the fairness level of our deployment. The MeGlass dataset (noted as a publicly available Kaggle release) is used as the base of experiments, which is an image-only face dataset to test the strength against eyeglasses, and the MeGlass dataset is a base to assess the reality of deployment through the Kaggle publicly available release. Using eyeglass wearers as a sample visually complex group, we conduct a label-free Responsible AI audit in the conditions when we do not have labels of the sensitive attributes. We observed that average performance difference in mean is not statistically significant (p = 0.36) by the permutation test, while mean difference in p99 is 16.85 ms between cohorts, bringing up fairness differences not captured by averages. Doing systems-level regression, we attribute such variances to pipeline-level attributes (e.g. detection confidence, input complexity) as opposed to the use of these attributes directly as causal factors. We also evaluate mitigation policies and determine trade-offs among the latency, fairness and recognition fidelity. The findings here show that fairness inequities can arise at worst-case runtime behavior, not in average latency, suggesting that the use of metrics focused on the tail of a run-time fairness distribution should be considered in fairness assessments of run-time deployed face recognition systems.
Короткий адрес: https://sciup.org/15020769
IDS: 15020769 | DOI: 10.5815/ijeme.2026.05.04