A Lightweight Face Anti-Spoofing Framework with Spoof Artifact Enhancement and Adaptive Feature Fusion

Mudunuru Suneel Banothu Yedukondala Venkata Naga Raja Swamy Madhava Rao Maganti Seva Sreedhar Babu P. Rama Koteswara Rao Vijaya Kumari Devarapalli Kama Ramudu

Журнал: International Journal of Engineering and Manufacturing @ijem

Статья в выпуске: 5 vol.16, 2026 года.

Бесплатный доступ

Face recognition is widely used for biometric authentication in applications such as mobile devices, financial services, intelligent surveillance, access control, and border security. However, face recognition systems remain vulnerable to presentation attacks, including printed photographs, replay attacks, and three-dimensional masks. Although recent deep learning-based face anti-spoofing (FAS) methods have achieved substantial improvements, many existing approaches still involve considerable computational cost, provide limited emphasis on fine-grained spoof artifacts, and face challenges in effectively integrating heterogeneous representations. To address these limitations, this paper proposes a Lightweight Face Anti-Spoofing Framework with Multi-Modal Representation, Spoof Artifact Enhancement, and Adaptive Feature Fusion. The framework starts from a single RGB facial image and constructs complementary Depth and Near-Infrared (NIR) representations using a Depth and Near-Infrared Construction Module (DNCM), rather than requiring dedicated Depth or NIR sensors. A shared EfficientNetV2 backbone is then employed to extract features from the three representations with reduced computational redundancy. The proposed Spoof Artifact Enhancement Module (SAEM) emphasizes subtle spoof-specific visual cues, while Cross-Modal Consistency Learning (CMCL) reduces representation discrepancies across the constructed modalities. Subsequently, the Adaptive Feature Fusion Module (AFFM) dynamically weights the refined representations according to their discriminative contribution. Extensive experiments on CelebA-Spoof, CASIA-SURF, HQ-WMCA, and SiW-M demonstrate the effectiveness of the proposed framework. The framework achieves accuracies of 98.16%, 97.54%, 99.08%, and 97.18%, with corresponding ACER values of 2.87%, 4.36%, 2.03%, and 4.18%, respectively. Across five independent runs, statistical analysis further indicates consistent performance with significant improvements over the selected baseline. In addition, the framework requires only 9.8 million parameters and 2.1 GFLOPs and achieves an inference speed of 82 FPS, demonstrating a favourable balance between detection effectiveness and computational efficiency. The results indicate that multi-modal representation combined with explicit spoof artifact enhancement and adaptive feature fusion can provide an efficient solution for robust and real-time face anti-spoofing.

Biometric Security \ Face Anti-Spoofing \ Presentation Attack Detection \ Multi-Modal Representation \ Spoof Artifact Enhancement \ EfficientNetV2 \ Cross-Modal Consistency Learning \ Adaptive Feature Fusion

Короткий адрес: https://sciup.org/15020731

IDS: 15020731   |   DOI: 10.5815/ijem.2026.05.21