GenAI-Driven Interview Performance Assessment: Revolutionizing Recruitment with AI Insights

Автор: Sathvik Vadarevu, Manasa Viriyala, Garlapati K.V.S. Sai Komal, Venkat Vinukonda, Jeethu V. Devasia

Журнал: International Journal of Information Engineering and Electronic Business @ijieeb

Статья в выпуске: 4 vol.18, 2026 года.

Бесплатный доступ

This paper proposes IPAMS (Interview Performance Assessment using Gen AI), which is an AI-driven platform that automates interview evaluations using advanced technologies like Convolutional Neural Networks (CNN) to gain insights on facial emotions and expressions, Large Language Models (LLM) to generate and process interview questions, YOLO (You Only Look Once) for real-time object detection, and APIs for speech-to-text transcription and behavioral analysis. The system captures video responses and analyzes key elements such as sentiment, speech patterns, body posture, and facial expressions, generating a detailed report. This report highlights a candidate’s strengths and areas of improvement and is sent directly to their email with actionable insights. IPAMS modernizes recruitment by providing unbiased assessments, saving time and resources for recruiters. For candidates, it offers a valuable mock interview tool, delivering feedback on technical skills, confidence, stress levels, and nonverbal communication. By combining cutting-edge AI and analytics, IPAMS delivers an efficient, objective, and insightful solution for recruitment and self-assessment, benefiting all stakeholders in the interview process.

AI-driven platform, Convolutional Neural Networks, Large Language Models, YOLO, GenAI

Короткий адрес: https://sciup.org/15020605

IDR: 15020605   |   DOI: 10.5815/ijieeb.2026.04.06

Текст научной статьи GenAI-Driven Interview Performance Assessment: Revolutionizing Recruitment with AI Insights

Published Online on August 8, 2026 by MECS Press

Interviews play a crucial role in the hiring process, serving as a vital tool for organizations to assess potential

This work is open access and licensed under the Creative Commons CC BY 4.0 License.

candidates and for individuals to showcase their skills and qualifications. However, traditional interview methods are often time-consuming, subjective, and prone to human bias [1, 2]. These inefficiencies can hinder the recruitment process, making it difficult for organizations to consistently find the right talent and for candidates to receive valuable feedback on their performance [3]. In addition, the pressure and stress that candidates face during interviews can affect their performance, which can lead to an incomplete or inaccurate assessment of their true capabilities [4, 5]. AI-powered interview automation has emerged as a game-changing solution to address these challenges. By integrating technologies such as Natural Language Processing (NLP), Machine Learning (ML), and sentiment analysis, AI can standardize and streamline the evaluation process, ensuring more objective and accurate assessments [2, 6]. These technologies allow AI systems to evaluate a candidate’s responses, identify key skills, and detect emotional cues such as confidence and stress, which may be difficult for human interviewers to assess consistently [7, 8]. AI-driven systems can assess not just technical skills but also behavioral traits like confidence, stress management, and body language, providing deeper insights into a candidate’s overall fit for the role [9–11]. As AI technology continues to evolve, the future of interview automation looks promising. Organizations are increasingly adopting AI-driven tools to enhance their recruitment strategies, while candidates benefit from more personalized feedback and better opportunities to showcase their abilities [2, 5, 6]. With AI, interviews can become a more objective, insightful, and effective means of assessing candidates, ultimately leading to better hiring decisions and improved organizational performance [1, 12].

The proposed system, IPAMS is designed to augment human recruiters by providing objective data, not to replace the essential role of human judgment in assessing a candidate’s overall suitability. It integrates advanced artificial intelligence technologies, such as natural language processing, computer vision, and speech analytics, to deliver a comprehensive and automated platform for evaluating interview performance. The system incorporates features including intelligent question generation, video-based response analysis, emotion recognition, confidence and speech pattern assessment, body language evaluation, and seamless resume integration. Collectively, these components enable a holistic and data-driven evaluation of a candidate’s interview behavior.

2.    Literature Survey

A significant base paper that inspired the development of IPAMS is “AI-Based Mock Interview Evaluator: An Emotion and Confidence Classifier Model” [6]. This research introduces an AI-driven mock interview evaluator designed to assess candidates on emotions, confidence, and knowledge. The system simulates real interview scenarios to help candidates refine their skills. The approach incorporates five phases, employing CNN for facial expression recognition [13, 14], NLP for speech analysis [7, 8], and web scraping for knowledge mapping. With a training accuracy of 80.75% and a testing accuracy of 76.34%, this system demonstrates the efficacy of integrating AI technologies for a holistic evaluation [6]. Future enhancements proposed include improving accuracy and extending its scope to diverse interview scenarios using advanced machine learning methods and larger datasets [15].

Another influential study is “A Comprehensive Study and Implementation of the Mock Interview Simulator with AI and Pose-Based Interaction” [4]. This research presents a mock interview simulator that combines AI and poses detection to create an immersive and realistic interview experience. The simulator uses AI-driven chatbots powered by GPT-3.5 Turbo for generating questions and feedback, speech recognition for analyzing user responses, and the Mediapipe framework for posture detection, providing feedback on body language [2, 4]. The modular architecture allows seamless interaction between various components, delivering accurate feedback on both verbal and non-verbal communication. Results from user testing indicated positive reception, highlighting the system’s effectiveness in providing a realistic and insightful mock interview experience. Future enhancements include integrating virtual reality for deeper immersion, more diverse question sets, and advanced body language analysis [11].

Another major reference is “A Semantic Approach for Automated Hiring using Artificial Intelligence & Computer Vision” [1]. This research paper introduces an automated interview system that leverages NLP and deep learning methods for conducting job interviews. The system is composed of four modules, with the first being the resume classification module, which is built on advanced NLP frameworks and the transformer-based BERT model. This model employs named entity recognition and text classification to extract key information from resumes.

The above studies emphasize the importance of AI in transforming the interview process by providing candidates with actionable insights and realistic feedback [1, 16]. These foundational concepts have been extended in IPAMS, leveraging advancements in CNNs [14, 17], LLMs [3, 15], YOLO for real-time object detection [10, 18], and APIs for speech-to-text transcription [19]. Unlike the referenced studies, IPAMS integrates additional layers of behavioral analysis, including stress detection, confidence evaluation, and technical proficiency assessment [2, 5, 6]. This holistic approach aims to revolutionize both mock interviews for candidates and the recruitment processes for organizations by offering a data-driven, unbiased, and efficient evaluation platform [1, 5, 12].

3.    Experimental Methods and Procedures

This section elaborates on the system architecture and underlying methodology of IPAMS.

  • 3.1.    Architecture of IPAMS

Fig. 1 illustrates the system architecture of this proposed method, focusing on the website architecture designed for scheduling and conducting interviews for candidates.

Fig. 1. System Architecture for Candidate Registration and Interview Scheduling.

The architecture begins with candidate registration, which is streamlined through a user-friendly website integrated with Google Forms. This integration allows for efficient data collection while maintaining a straightforward and accessible experience for candidates. The registration process captures a wide range of essential details, including the candidate’s name, email address, and phone number, ensuring accurate identification and communication. Candidates are also asked to specify their preferred interview type (HR or Technical) and their preferred interview mode (online or in-person), accommodating flexibility based on their needs. Additionally, the form provides fields to record any accessibility needs or special requests, ensuring an inclusive experience tailored to individual requirements.

Once candidates submit their information through the embedded Google Form, the system organizes the data into structured formats that are readily available for further processing. This centralized approach minimizes manual errors, enabling seamless handling of multiple candidates. After completing the registration, candidates are directed to schedule their interview by selecting a convenient date and time through the same platform. The Google Form link also includes an option for candidates to upload their resumes, ensuring all critical documents are collected in one step. This dual functionality eliminates the need for multiple touchpoints and keeps the process efficient and streamlined.

Following the completion of registration and scheduling, the system dynamically generates a unique interview link for each candidate. This link acts as a gateway to their personalized interview session, ensuring a smooth and private experience. The link, along with detailed information about the scheduled interview - including the chosen interview type, mode, and timing - is shared with the candidate via an automated email notification. The email communication is designed to be clear and informative, reducing the likelihood of miscommunication or missed sessions.

On the day of the interview, candidates can effortlessly use the provided link to access their session. This setup not only simplifies the candidate’s journey but also ensures that all necessary steps - from registration to participation—are handled systematically and without complications. The integration of Google Form, email notifications, and scheduling into a cohesive process underscores the architecture’s focus on automation, efficiency, and user-centric design.

  • 3.2.    Working Methodology

  • 3.3.    CNN Architecture for Facial Expression Recognition

    The proposed CNN architecture is designed to analyze video frames and extract facial expression features. This architecture is trained on the FER-2013 dataset [21], which contains a large number of labeled facial images. The CNN architecture typically consists of multiple convolutional layers, pooling layers, and fully connected layers, as given in

The IPAMS system is an innovative platform that automates and enhances the candidate evaluation process, integrating advanced technologies for a seamless and comprehensive experience. Its working methodology is depicted in Fig. 2.

Fig. 2. Overview of IPAMS System Workflow.

The process starts with candidates registering on the platform through a user-friendly interface, where they provide essential details such as name, email, phone number, and preferred interview type (HR or Technical). They also upload their resumes directly on the platform. The uploaded resume is processed using the PyPDF2 library to extract detailed text data, including the candidate’s educational qualifications, skills, and work experience. This extracted information is structured and sent to a powerful LLM, Gemini [20] along with a carefully designed prompt to generate personalized interview questions tailored to the candidate’s expertise and selected interview type. These questions encompass a range of topics, including technical knowledge, problem-solving abilities, and behavioral insights, and are stored in the system for real-time display during the interview session. The extracted text from the resume is used by the LLM to generate questions. During the interview, the video and audio are recorded. The audio is transcribed to text using AssemblyAI, and this text, along with the video, is input into the LLM for analysis. Simultaneously, the video is processed by the CNN to detect facial expressions. The outputs from the LLM and the CNN are combined to create the final report.

Candidates answer these questions one at a time by recording their responses through a live video recording feature integrated into the platform, ensuring an interactive and dynamic assessment. The recorded videos are uploaded and processed by advanced APIs such as AssemblyAI, which convert speech into text with high accuracy. The transcriptions are then analyzed to evaluate the candidate’s confidence levels using tone analysis, while the text content undergoes natural language processing to assess the depth and relevance of their responses. Additionally, emotion detection models analyze facial expressions, identifying cues such as happiness, nervousness, or frustration, while posture analysis algorithms evaluate body language to gauge the candidate’s composure and engagement. LLMs are selected for their ability to generate contextually relevant questions and analyze complex textual data from candidate responses. YOLO is chosen for real-time object detection.

To ensure integrity, the system employs sophisticated malpractice detection algorithms to identify inconsistencies or suspicious behaviors, such as unusual movements or off-screen assistance. All these analytical outputs, including speech-to-text transcriptions, emotion scores, confidence ratings, and behavioral assessments, are compiled and sent back to the LLM for synthesis. The LLM generates a detailed evaluation report that provides an in-depth analysis of the candidate’s strengths, weaknesses, and overall performance. This report also includes constructive feedback and suggestions for improvement, enabling candidates to refine their interview skills. The final evaluation report, formatted as a professional PDF document, is appended with all the generated insights and automatically emailed to the candidate, ensuring prompt delivery and a streamlined experience. This fully automated, AI-driven system redefines the interview process, making it objective, efficient, and insightful for both candidates and recruiters.

responses lacked the depth compared to more specialized models. BERT, with its fast responses, performed excellently in classification tasks but struggled with long, coherent generative responses. This made it suitable for specific tasks such as categorizing or classifying interview responses but less effective for generating comprehensive insights. Finally, Alpaca, while fast and effective at following instructions, was limited in its analytical depth. Its responses were generally effective for straightforward tasks but fell short when a deeper analysis was needed. Overall, Gemini and Llama provided the best balance of context understanding, response quality, and analytical depth, proving to be the most effective models for generating comprehensive interview evaluations in IPAMS.

Table 1. Software Libraries and Their Purposes in the IPAMS Project.

Library

General Purpose

Purpose in IPAMS

os

Provides a way of using operating system-dependent functionality

Set and manage environment variables for API keys

Streamlit

A web application framework for building interactive apps

Create the user interface for the interview analysis tool

Streamlit.components.v1

Enables embedding custom HTML/CSS components in Streamlit apps

Add video recording functionality for candidates' answers

Google.generativeai

API for Google’s generative AI capabilities

Generate interview questions and analyze answers using AI

Assemblyai

Speech recognition API for transcribing audio

Transcribe video answers into text for analysis

Pypdf

Library for reading and manipulating PDF files

Extract text from candidates’ resumes in PDF format

Fpdf

Library for generating PDF documents

Create and save the interview analysis report as a PDF

Smtplib

A built-in Python library for sending emails

Send the generated PDF report to the candidate's email

Email.mime.text

Part of the email package for creating email messages

Create the body of the email to accompany the PDF attachment

Email.mime.multipart

Part of the email package for handling multipart email messages

Construct the email with the PDF attachment

Email.mime.application

Part of the email package for including files in emails

Attach the PDF report to the email

Langchain

A framework for building applications with LLMs

Facilitate the integration of large language models in generating and analyzing interview questions and responses

Tenserflow

It is widely used for building and training various deep learning models, including neural networks

Used for building a CNN algorithm for confidence and stress prediction

It is important to acknowledge the limitations of LLMs. ChatGPT sometimes produced generic responses [22], Gemini has higher computational costs [20], Llama occasionally lacked depth [23], BERT struggled with longer responses [24], and Alpaca had limited analytical depth [25]. To mitigate these limitations, IPAMS combines LLM analysis with other data sources, such as facial expression and speech analysis. A comparison of these LLM models is given in Table 2.

Table 2. Comparison of LLM Models in IPAMS.

Model

Response Time

Response Quality

Analysis Quality

ChatGPT [22]

Moderate

High coherence and contextually relevant

Good for generating insights and summaries

Gemini [20]

Moderate

Strong contextual understanding, varied responses

High analytical capabilities for detailed analysis

Google Flash

Fast

Effective and context-aware responses

Good for quick insights but less depth in analysis

Llama [26]

Moderate

Strong understanding of context, high versatility

Good performance in generating accurate analyses

BERT [27]

Fast

Excellent for classification tasks, less coherent in longer responses

Limited analytical capabilities in generative contexts

Alpac a [25]

Fast

Effective at instruction following, competitive

Limited analytical depth compared to larger models

  • 4.2.    IPAMS-Developed Features

  • 4.3.    Facial Expression Detection using CNN

    The IPAMS system utilizes the FER-2013 dataset, a widely recognized facial expression recognition dataset, to train the CNN for predicting confidence and stress levels during the interview. The FER-2013 dataset contains labeled facial images annotated with various emotional states such as happiness, sadness, anger, and surprise, which are used to train the model to accurately recognize and classify these emotions. The CNN architecture processes these images through multiple convolutional layers, followed by max-pooling layers to reduce dimensionality. Techniques like dropout and batch normalization were applied to prevent overfitting and improve generalization. The model concludes with fully connected layers that classify emotional states based on facial expressions.

  • 4.4.    Quantitative Evaluation

    The CNN model has achieved an average accuracy of 65% in the FER-2013 test set. The LLM’s performance is evaluated by comparing its generated interview question with a set of human-generated questions, achieving a relevance score of 0.85 (on a scale of 0 to 1). The overall system’s effectiveness is measured by correlating IPAMS scores with human interviewer ratings in a mock interview setting, resulting in a correlation coefficient of 0.78.

Our proposed model, the AI-powered Interview Performance Assessment System, integrates advanced NLP, video analysis, and machine learning techniques to revolutionize candidate evaluation. Combining emotion, confidence, knowledge, and body language analysis, IPAMS offers one of the most comprehensive approaches in automated interview evaluation. The emotion analysis module uses a CNN-based model to detect facial expressions and emotional cues, providing insights into stress, confidence, and other emotions critical to performance. The confidence analysis system utilizes AssemblyAI for real-time speech-to-text transcription, assessing vocal features like pitch and speaking rate to evaluate confidence levels effectively. For knowledge assessment, Llama 2 generates tailored interview questions based on a candidate's resume, ensuring relevance. Responses are analyzed for clarity, depth, and accuracy. Body language analysis, using CNNs, evaluates posture, head pose, and non-verbal cues, offering a holistic view of the candidate’s demeanor.

The system personalizes interviews through dynamic question generation and resume-based technical queries. Personality trait analysis combines textual and visual data to provide insights into communication style and emotional intelligence. Malpractice detection ensures integrity by analyzing gaze and non-verbal cues for inconsistencies. IPAMS continuously evolves with data, enhancing its accuracy in emotion, confidence, and personality prediction. This AIdriven platform delivers reliable, insightful, and unbiased assessments, setting a new benchmark in interview evaluation. IPAMS differentiates itself by offering a more comprehensive analysis. For example, while [6] uses Mediapipe for emotion detection, IPAMS employs a CNN, enabling more detailed and accurate emotion recognition. Additionally, IPAMS integrates resume analysis directly into the question generation process, ensuring higher relevance than systems like [4] that use more generic question generation methods.

Table 3 gives a comparison of IPAMS with existing AI-based interview evaluation systems.

Table 3. Comparison of IPAMS with Existing AI-Based Interview Evaluation Systems.

Features

Reference Paper [6]

Reference Paper [4]

Reference Paper [1]

IPAMS

Main Idea

AI-based interview simulator with pose-based feedback

Automated hiring through AI and computer vision

AI-driven mock interview platform for performance analysis

Combines all features, improving them using advanced NLP and video analysis

Emotion Detection

Mediapipe for posture and emotion detection

CNN and LSTM for facial expressions

Analyzes eight emotional features from facial expressions

Emotion analysis using advanced CNN and ML techniques

Confidence and Speech Analysis

Speech recognition for transcriptions

LSTM for voice modulation analysis

Audio analysis for pitch, amplitude, and speaking rate

Uses AssemblyAI for realtime transcription and nuanced confidence analysis

Knowledge/ Response Analysis

GPT-3.5 Turbo for question generation and feedback

BERT for resume-job similarity, NLP for responses

Text-based analysis for DISC personality traits

Uses LlaMA 2 for tailored question generation based on resume and advanced response analysis

Body Language and Head Pose Detection

Posture detection using Mediapipe for real-time feedback

Malpractice detection using OpenCV

Analyzes head movements and angles for posture

IPAMS CNN model for posture and head pose, offering holistic body language analysis

Interview Customization

Customizable interview questions

Real-time question generation via NLP

Customizes questions based on job type

Dynamic question generation based on resume and job type

Resume Integration

None

Resume screening via BERT

None

Directly integrates resume analysis to generate technical questions and feedback

Speech Modulation

Speech analysis via MediaPipe

LSTM models for voice modulation

Analyzes pitch, speaking rate, and amplitude from audio

Advanced speech analysis including fluency, modulation, and emotion

Personality Trait Analysis

None

Personality analysis using AI and NLP

DISC and intrinsic personality traits

Provides personality analysis based on text and body language

Malpractice Detection

None

YOLO for malpractice detection

Monitors gaze and behavior for malpractice detection

Can integrate behavior Tracking and malpractice detection during interviews

Through experimentation with different configurations, including the number of layers, filter sizes, and optimization methods, significant improvements in the model’s performance were observed. The CNN successfully detected emotional cues related to confidence and stress from the FER-2013 dataset, contributing valuable insights to the overall assessment. This facial expression analysis, in conjunction with speech and body posture evaluations, enhanced IPAMS’ ability to assess both verbal and non-verbal cues, providing a comprehensive evaluation of the candidate’s psychological state during the interview. Its loss and accuracy are given in Figures 4 and 5, respectively.

Fig. 4. Training and Validation loss.

5. Conclusion and Future Scope

The Interview Performance Assessment using GenAI system integrates advanced technologies, including natural language processing, computer vision, and speech analysis, to offer a comprehensive, AI-powered platform for interview performance evaluation. By combining automated question generation, video response analysis, emotion detection, confidence and speech analysis, body language evaluation, and resume integration, IPAMS aims to provide an unbiased, detailed, and efficient assessment of a candidate’s interview performance. The system offers a more streamlined, objective approach to interview evaluations, allowing for real-time feedback that can help candidates identify strengths and weaknesses in their responses, body language, and emotional state. Through the integration of large language models such as Llama 2 for question generation and response analysis and utilizing state-of-the-art speech recognition and emotion detection technologies like AssemblyAI and advanced CNNs, the project has made significant strides in creating a platform that can simulate a real interview experience. It not only helps candidates prepare more effectively but also support recruiters by providing a more accurate and holistic view of candidate performance, far beyond the scope of traditional interview processes. The results obtained from the various AI models used in the project have shown promising accuracy and efficiency in areas such as emotion detection, confidence assessment, and response analysis. The combination of video, audio, and text analysis allows the platform to capture multidimensional aspects of the interview process, resulting in a richer evaluation than any single method alone. However, challenges remain, especially in areas such as integrating real-time feedback across multiple modalities, improving model accuracy, and ensuring that the system can scale effectively to handle diverse candidate profiles and interview scenarios.

Fig. 5. Training and Validation Accuracy.

Future work will include evaluating the system for potential biases between different demographic groups. This will involve analyzing model performance on diverse datasets and implementing bias mitigation techniques. Although IPAMS provides a detailed analysis of various candidate attributes, it is crucial to remember that human recruiters bring invaluable contextual understanding and empathy to the hiring process. IPAMS aims to support and enhance, not substitute, this human element. Future research will involve pilot studies with real recruiters to evaluate the integration of IPAMS into their workflow and assess its impact on the efficiency and effectiveness of the recruitment process. The system could be enhanced to provide candidates with feedback and resources on stress management techniques. Furthermore, the evaluation algorithm could be refined to account for the potential influence of anxiety on performance, ensuring a fairer assessment. Future development will prioritize user experience (UX) enhancements [28]. This includes streamlining the interface for both candidates and recruiters, providing clearer instructions, and improving the accessibility of evaluation reports.

All the Declarations and StatementsAuthor Contributions Statement

Sathvik Vadarevu – Conceptualization, Methodology, and Writing: Proposed research ideas, Constructed the overall framework, Drafted the initial manuscript, contributed to the literature survey, and documented the technical background of the study.

Manasa Viriyala – Data Curation and Software Implementation: Handled data acquisition, dataset preprocessing, and implementing the research model.

Garlapati K. V. S. Sai Komal – Model Training, Validation, and Performance Evaluation: Led the model training process, validated results using standard metrics, and benchmarked performance against existing methods.

Venkat Vinukonda – Formal Analysis, Visualization, and Statistical Analysis: Performed in-depth analysis of experimental results, prepared performance charts, and ensured the statistical robustness of the evaluation.

Jeethu V. Devasia – Supervision and Writing - Review and Editing: Supervised project execution, Reviewed and edited the manuscript, ensured clarity and coherence, and helped coordinate project milestones and deadlines.

All authors have read and agreed to the published version of the manuscript.

Conflict of Interest Statement

The authors declare no conflicts of interest.

Funding Declaration

None.

Data Availability Statement

The application developed for this study is publicly accessible at

Ethical Declarations

This study does not involve human or animal subjects; hence, ethical approval was not required.

Acknowledgments

We sincerely thank the experts for their professional evaluation and valuable recommendations, which have contributed to improving the quality of the experiment and the reliability of its results.

Declaration of Generative AI in Scholarly Writing

The authors acknowledge the use of OpenAI GPT-5 solely for language enhancement. All intellectual contributions, interpretations, and conclusions are those of the authors.

Abbreviations

The following abbreviations are used in this manuscript:

IPAMS - Interview Performance Assessment using

Gen AI - Temporal Convolutional Network

CNN - Convolutional Neural Networks

LLM - Large Language Models

YOLO - You Only Look Once

NLP - Natural Language Processing

ML - Machine Learning

UX - User Experience

APPENDIX A\B\C…, with appendix tile

None.