**Exploring Employee Working Productivity: Initial Insights from Machine Learning Predictive Analytics and Visualization**

\*\*This is a Double-blind review, please do not include authors information in this version \*\*

Received Date: \*date

Accepted Date: \*date

Published Date: \*date

**HIGHLIGHTS**

  - **Developed a predictive analytical model using machine learning to explore and predict employee working productivity within organizations.**

  - **Employed ranker algorithms to identify and assess the significance of attributes influencing employee working performance in the organizational context.**

  - **Utilized data visualization techniques for descriptive analytics, providing initial insights and correlations among various attributes of employee working patterns.**

  - **Implemented feature encoding in data preprocessing, ensuring relevant features for effective training of the employee working productivity prediction model.**

**ABSTRACT**

*Employee working productivity prediction is vital for effective resource allocation, increased productivity, and upholding a high-performance culture in organizations. However, predicting employee productivity and understanding the root factors influencing working performance pose significant challenges. Traditional human resource management practices often lack data-driven insights, resulting in poor resource allocation and productivity enhancement strategies. Subjective assessments, numerous assessment factors, and difficulties in interpreting predictive mechanisms add to the complexity of the task. To address these challenges, this research aims to develop a predictive model using machine learning techniques to determine employee productivity within organizations. Data from an academic institution were collected and pre-processed by encoding relevant features before applying various machine learning predictive models. Decision tree regressor, linear regression, MLP regressor, random forest regressor, SGD regressor, voting regressor, and Xgboost regressor were employed as predictive models. Ranker algorithms, including InfoGainAttributeEval, GainRatioAttributeEval, and CorrelationAttributeEval, were utilized to identify the most significant attributes affecting employee working performance. Additionally, descriptive analytics techniques were employed to visualize the data, extracting valuable insights and understanding the correlations among the features. Experimental results revealed that the linear regression model achieved the best performance in terms of Mean Absolute Error (MAE) and Mean Squared Error (MSE), with values of 0.4878 and 0.4682, respectively. Thus, the linear regression model emerged as the most accurate predictor for employee productivity in the given organizational context. Based on these findings, it is recommended that organizations consider adopting linear regression for predicting employee productivity. The research findings also highlighted certain attributes that play an imperative role in predicting employee performance. Attributes such as "Department," "Actual Productive hours," "Internet Speed," and "COVID-19 adoption month" emerged as highly influential factors across multiple ranking techniques. The data visualization provided valuable insights into various aspects of employee performance, such as productivity trends before and after the pandemic, departmental performance, internet connectivity's impact on productivity, age-related trends, overtime distribution, and promotion rates. Organizations can use this data to inform workforce planning, address specific challenges in departments, and cultivate an inclusive work environment. By regularly assessing productivity data and implementing recommended strategies, organizations can enhance productivity, create a conducive work environment, and support employee well-being and growth. Future research can explore more advanced machine learning algorithms, incorporate time-series analysis for temporal dependencies, and expand data collection from diverse organizational settings to improve the generalizability of predictive models.*

***Keywords:** Employee Productivity; machine learning; prediction; visualization*

# INTRODUCTION 

The escalation of costs within a company or organization, such as meeting minimum salary requirements, significantly influences how the organization manages its employees. Research indicates that when individuals do not align with the company's culture or values, they are often encouraged to voluntarily leave rather than facing termination. Conversely, termination can present challenges in securing future employment opportunities, making it crucial for organizations to strike a balance between cost management and employee well-being. Thus, an effective evaluation of the employee working performance system is an important factor for workforce management such as sustaining working productivity, optimal resource allocation and employment-related decision making. Inadequate workforce management practices can result in substantial losses for organizations (Sabuj et al., 2023). The current practices of the employee evaluation performance system possess several challenges due to assessment’s subjectivity, assessment factors variety, and assessment inconsistencies (Sarker et al., 2018). Furthermore, lack of data-driven approach to predict and improve employee working productivity in organizations. Existing methods may be limited in their ability to capture complex patterns and interactions among various attributes that affect productivity in more specific context or organization (Hu, 2021; Li et al., 2021). Additionally, the absence of effective data visualization techniques hinders the ability to gain valuable insights from data. With the increasing uncertainty and flexibility in the working environment, it becomes even more crucial to analyse employee working performance and identify the factors that affect productivity.

Machine learning techniques were applied to build predictive models that can accurately forecast employee working productivity. By analyzing vast amounts of data related to various attributes like working hours, department, internet speed, and more, machine learning algorithms can identify patterns and correlations that may not be apparent through traditional methods. Data visualization complements machine learning by presenting complex data in a visually intuitive manner. Organizations can gain a deeper understanding of the factors influencing employee productivity. Visualizing the relationships between different attributes allows organizations to spot trends, identify bottlenecks, and explore opportunities for improvement.

Therefore, the primary objective of this research is to develop a predictive model using machine techniques to determine employee productivity for one of the academic institutions in Sarawak, Malaysia. The data related to the employee working patterns, working hours and working preferences were collected, cleaned and pre-processed by using feature encoding techniques for further processing. Various machine learning predictive models, including decision tree regressor, linear regression, MLP regressor, random forest regressor, SGD regressor, voting regressor, and Xgboost regressor, were applied to predict employee productivity. Additionally, ranker algorithms such as InfoGainAttributeEval, GainRatioAttributeEval, and CorrelationAttributeEval were employed to identify the most significant attributes affecting employee working performance. Descriptive analytics techniques by using Tableau software were utilized to visualize, plot, and analyze the data, extracting valuable insights and understanding the correlations among the attributes. By developing predictive model, the institution can optimize their resource allocation, identify areas for improvement, and create a productive work environment. In addition to that, the use of ranker algorithms offers valuable insights into the underlying factors that influence employee productivity within the organization.

This paper is organized as follows: first, we present a literature review, followed by the methodology used in the study. Next, we present the experimental results obtained from our analysis, and finally, we draw conclusions and provide recommendations based on our findings.

**LITERATURE REVIEW**

In recent years, the application of machine learning algorithms in various fields has gained significant attention due to their potential to provide valuable insights and predictions. In this literature review, we explored several studies that utilize machine learning techniques to predict and evaluate employee performance, turnover intention, fatigue, and churn within different organizational contexts.

The research conducted by Sarker et al., (2018) focuses on developing an effective employee performance evaluation system and predicting future employee performance using K-Means Clustering and Decision Tree Algorithm. The results from the prediction may aid decision-making processes within an organization, predicting employee performance for the next year, identifying inefficient employees, and facilitating decision-making for promotions or designations. The study takes into account factors such as personality, punctuality, tact, oral expression, and other performance evaluation criteria to evaluate employee performance. The research proposed a hybrid procedure that combines Data Clustering and Decision Tree algorithms employee performance prediction. The results of the study demonstrate the successful application of K-Means Clustering and Decision Tree algorithms for employee performance evaluation and prediction. The findings indicate that clustering can be employed to group employees based on their characteristics and abilities, while decision tree algorithms assist in making decisions related to employee development and qualifications. Future works are suggested to gather more data from different companies and apply the algorithms to predict employee performance in various organizations. This could potentially enhance the accuracy and applicability of the developed performance evaluation system.

In the study conducted by Tambde & Motwani (2019), they focused on the issue of employee churn and its impact on company performance. An expert prediction system by using machine learning algorithms was developed to forecast the rate of employee attrition, thereby reducing costs and improving employee retention to foster company growth. The factors that influence churn have been identified and a predictive model was created based on employee working details and situations. Various machine learning algorithms, including Random Forest, were employed for churn prediction and has provided insights into potential churn risks. Performance metrics, such as accuracy and error rate, were used to evaluate the effectiveness of the machine learning algorithms. The study found that Random Forest outperformed other classifiers in predicting employee churn, showcasing its suitability for this application. Other than that, the prediction model is also can be used in identifying employees likely to leave, enabling organizations to take proactive measures to retain their workforce effectively. However, the study acknowledged limitations related to overfitting or underfitting issues in machine learning algorithms. These limitations could potentially impact the accuracy of the churn prediction model and warrant further attention in future research. Additionally, they proposed exploring additional data sources to augment the churn prediction capabilities, providing deeper insights into the factors influencing employee attrition.

The study conducted by Jayadi et al. (2019) predicting employee performance and turnover using machine learning techniques by using Naïve Bayes classification algorithm that aimed to reduce costs and promote company growth through better workforce management. The prediction model provides valuable assistance to human resource departments in predicting and managing employee turnover and performance based on key performance indicators (KPIs), employee satisfaction, and other factors that influence performance and turnover. The Naïve Bayes classification algorithm, along with CRIPS-DM methodology, was employed in predictive modelling. The results demonstrated that the Naïve Bayes algorithm achieved high accuracy (95.48%) in predicting employee performance. Future works were proposed to further improve the prediction accuracy. This could involve fine-tuning the model and exploring additional factors that could enhance the accuracy of employee performance prediction.

A systematic literature review to address the fragmented understanding of factors affecting employee performance was carried out by Atatsi et al. (2019). The study aimed to synthesize existing literature on organizational citizenship behavior, leader-member exchange, learning, innovative work behavior, and employee performance across various countries, disciplines, and organizations, with a specific focus on identifying knowledge gaps in the African context. The prediction variables in this study included organizational citizenship behavior, leader-member exchange, learning, innovative work behavior, and employee performance. The study's findings revealed positive relationships between the identified behaviors (organizational citizenship behavior, leader-member exchange, learning, innovative work behavior) and employee performance. The research also highlighted the significance of context and culture in understanding these relationships, particularly in the African context.

Monisaa Tharani & Vivek Raj (2020) addressed the issue of employee turnover intention in the IT & ITeS industry, which can significantly impact organizational performance and profitability. This study predicts turnover intention using machine learning algorithms and identify the significant factors influencing turnover intention in this industry to help them in developing strategies and policies to reduce turnover intention which improving employee satisfaction and engagement. The prediction model trained based on the attributes of employee turnover intention, alternative job opportunity, gender, education, willingness to relocate, job stress, attitude towards COVID-19, and other relevant factors. XG Boost algorithm and Logistic Regression were employed for predictive modelling. The results of the study indicated that the XG Boost algorithm achieved a high accuracy rate of approximately 94%. Logistic Regression identified several influencing variables, including alternative job opportunity, gender, education, willingness to relocate, job stress, attitude towards COVID, dependent family members, job role, job satisfaction, organizational commitment, and marital status. However, the study had some limitations, such as data collection based on turnover intention, which could change over time due to dynamic factors. Additionally, the research did not focus on employees who had already left the company, potentially leaving out critical insights. Furthermore, the data collected might not fully represent the entire IT & ITeS industry. For future works, the study suggested focusing on employees who have switched organizations to gain a broader perspective on turnover intention. Moreover, similar studies could be conducted in other sectors to predict turnover intention in different industries.

Li et al., (2021) created a prediction model to identify high-performing employees in Company A which is a publishing house and technology hub. The prediction model is used to assist the company in evaluating and understanding its internal operations and improve employee performance by analyzing various employee attributes and features. Various supervised classifiers, including Logistic Regression, Decision Tree, and Naïve Bayes, were utilized for predictive modeling of employee performance. The study compared the performance of these classifiers using Microsoft Azure Machine Learning and identified significant variables in the prediction model through Permutation Feature Importance. The target variable was Performance Rating, categorized into Excellent and Good, while predictive variables included gender, tenure group, marital status, business title, department, compensation grade profile, address city, program cost/head, service provider, training title, and type. The Two-Class Logistic Regression model exhibited the highest precision among the classifiers. Notably, features like Compensation Grade Profile, Tenure Group, and Service Department were found to be significant in the model. The researchers suggested that incorporating additional features could further enhance the prediction model's performance.

Saputra & Purwitasari (2022) conducted a study focusing on Fatigue Management in the mining industry. to address the issue of mining workers experiencing fatigue, which can have adverse effects on their performance and may lead to work-related incidents. The prediction variables such as sleep patterns, drug consumption, and other factors related to fatigue were used to develop the fatigue prediction model. The Random Forest algorithm was employed for predictive modelling and has outperformed Decision Tree and Logistic Regression, achieving an accuracy of 95.4%. Data for the study were collected from mining employees and categorized into fit and unfit groups and used the SMOTE technique to balance the dataset.

The findings indicated that sleep patterns and drug consumption played a major role in predicting employee fatigue. The model could be used as an early warning system to improve employee performance and reduce accidents. The limitations of this study including the use of survey data instead of real-time data from sensors or medical history comparisons. Future improvements could consider incorporating real-time data for more accurate predictions. Additionally, more data balancing techniques may be required for imbalanced datasets.

In the study conducted by Sabuj et al. (2023), they noticed that workers in this industry sometimes did not meet their productivity targets, causing losses for businesses. To improve productivity, they used machine learning to predict how well garment workers would perform. The data was collected from a garment production company in Bangladesh and analyzed factors related to worker performance. They used various machine learning techniques to develop the model, and the one that worked best was the Random Forest algorithm. One exciting finding was that the model revealed the sleep patterns and drug consumption, influenced worker productivity the most. This information could help managers make better decisions and improve worker performance. While the study focused on one company in Bangladesh, they believed their model could be useful for other companies too. They suggested further research to test its effectiveness across different businesses.

Hu (2021) used machine learning in predicting absenteeism at work in a courier company in Brazil, as absenteeism causing decreased productivity and financial losses for companies and leaders often rely on subjective judgment. By using machine learning, more objective and efficient prediction of employees' absenteeism can be demonstrated as well as helping leaders adjust working methods to minimize its impact on company finances and morale. The absenteeism determined using personal information of employees and reasons for absence. Four machine learning models, including linear models, decision trees, Random Forest, and AdaBoost, were employed to predict absenteeism. The integrated learning approach using Random Forest and AdaBoost achieved the highest accuracy of around 0.64. Descriptive statistical analysis revealed variables with clear correlations to absentee time, such as age, educational degree, and season. However, the integrated learning approach sacrificed interpretability and intuitiveness for improved accuracy. The study suggested future works to collect more data from wider resources and explore additional related factors to enhance the prediction of absentee time. By obtaining more comprehensive data and information about employees from different sources worldwide, a more accurate prediction of absentee time could be achieved.

Naz et al. (2022) conducted a study on predictive modelling of employee churn analysis for the IoT-enabled software industry. The predictive model is necessary to assess and identify potential churners as the employee churn in this industry incurs additional costs for organizations. The study analysed factors such as workload, working conditions, salary, job experience, and job satisfaction level to predict employee churn. The predictive model was built by utilizing four Exploratory Data Analysis (EDA) techniques and five Machine Learning (ML) algorithms. The Chi-squared-based decision tree algorithm achieved an impressive accuracy of 98%. The key factors affecting employee churn have been identified including satisfaction level, number of projects, time spent in the company, and last evaluation. Comparison with the RFE-based model showed improved results with the Chi-squared-based model. However, the limitations of this study using only filter-based methods and future work may involve exploring hybrid methods combining filter and embedded-based methods for more accurate and comprehensive employee churn prediction models.

In conclusion, the literature highlights the potential of machine learning algorithms in improving employee performance evaluation, prediction, and churn analysis. However, research problems exist in terms of model validation, overfitting or underfitting issues in dataset, exploring additional factors for prediction, considering cultural contexts, and enhancing data collection methods for more accurate and applicable predictive models. Future research should focus on addressing these issues to maximize full potential of machine learning in workforce management and decision-making processes within organizations.

# METHODOLOGY 

The methodology section presents an overview of the experimental setup and data analysis techniques employed to investigate the relationship between various factors and employee productivity in our research. This section outlines the key components of our research methodology, including the experimental setup, dataset preparation, feature ranking algorithms, machine learning prediction algorithms, and data description analytics and visualization techniques.

**Experimental Setup**

The data acquisition, experiments, data plotting and data visualization were carried out by using the programming environment, libraries and tools using Python libraries based on Spyder 4.2.2 and Tableau Desktop Professional Edition. Specifically, the feature encoding and classifications were performed by using Scikit-learn and Keras libraries.

**Dataset Preparation**

The dataset was collected from one of the higher education institutions in Sarawak. The data has 16 attributes associated with the employee working information and 103 number of instances. The details of the dataset are shown in Table 1.

**Table** **1.** Dataset Description

| No | Attributes                | Description                                                                           | Data Type  |
| -- | ------------------------- | ------------------------------------------------------------------------------------- | ---------- |
| 1  | Gender                    | Sexuality of employee (Male/Female).                                                  | Nominal    |
| 2  | Age                       | Age of employee.                                                                      | Continuous |
| 3  | Department                | Department name of employee.                                                          | Nominal    |
| 4  | Total Staff               | Number of staff in each department.                                                   | Discrete   |
| 5  | Working Promotion         | Employee’s working promotion (Yes/No)                                                 | Ordinal    |
| 6  | Daily Working Hours       | Total hours the employee worked every day.                                            | Discrete   |
| 7  | Monthly Working Days      | Total days the employee worked every month                                            | Discrete   |
| 8  | Overtime hours            | Total of overtime for employee in a day.                                              | Discrete   |
| 9  | Idle hours                | Total in hours the employee does not worked during working time every day.            | Discrete   |
| 10 | Targeted Productive Hours | Total hours the employee targeted to be productive in a day.                          | Discrete   |
| 11 | Actual Productive hours   | The total hours of employee really productive in a day.                               | Discrete   |
| 12 | Internet Speed            | The level of internet performance in the workplace.                                   | Ordinal    |
| 13 | Internet Problem Solution | The alternatives chosen to deal with Internet connectivity problems.                  | Nominal    |
| 14 | WFH Challenges            | The challenged faced during working from home.                                        | Nominal    |
| 15 | COVID adoption month      | The total of month the employee required to adopt the new norms of pandemic Covid-19. | Discrete   |
| 16 | Working Preference        | The types of working preferred by the employee.                                       | Nominal    |

**Feature Ranking Algorithms**

Feature ranking plays a role in identifying the most relevant and informative attributes within a dataset. The goal of feature ranking in predicting employee working performance is to pinpoint the essential elements that have the biggest impact on a worker's performance. The organization can evaluate the relevance and influence of many attributes in predicting employee performance outcomes by prioritizing the features through ranking. In this study, several ranking algorithms were employed including InfoGainAttributeEval, GainRatioAttributeEval, CorrelationAttributeEval, OneRAttributeEval, and ReliefF. These algorithms provide different perspectives and metrics for ranking attribute such as information gain, gain ratio, correlation, simplicity, and discriminative power. By considering multiple ranking algorithms, organizations can gain a comprehensive understanding of the attributes that consistently emerge as influential across different methods. This justification highlights the importance of feature ranking in selecting the most impactful attributes for predicting employee performance and justifies the use of multiple ranking algorithms to ensure robustness and accuracy in the analysis.

The InfoGainAttributeEval algorithm (Alhaj et al., 2016) measures the information gain provided by each attribute, taking into account the reduction in entropy or uncertainty in the data when a particular attribute is known. GainRatioAttributeEval, on the other hand, builds upon information gain by normalizing the results to account for the intrinsic bias towards attributes with many distinct values.

CorrelationAttributeEval (Billson et al., 2021) assesses the strength and direction of the linear relationship between each attribute and the target variable. It quantifies how well an attribute's values can be predicted using linear regression based on the target variable. OneRAttributeEval (Patil, 2014) evaluates the attribute's usefulness by constructing simple "one-rule" classifiers based on each attribute individually. It measures the accuracy and simplicity of the rules generated by each attribute. Lastly, ReliefF (Relief Feature Selection (Zainudin et al., 2018) algorithm evaluates the relevance of attributes by measuring the ability to distinguish between instances of different classes. It considers the nearest neighbors of instances to assess the weight or importance of each attribute.

**Machine Learning Prediction Algorithms**

The machine learning prediction algorithms provide insights and predictions that can assist organizations in making decisions related to workforce planning, talent management, and performance improvement strategies. In this study, there are various machine learning prediction algorithms have been evaluated for creating a predictive model of the employee working performance. The algorithms chosen for evaluation include Linear Regression, Random Forest Regressor, XGBoost Regressor, Decision Tree Regressor, MLP Regressor, SGD Regressor, and Voting Regressor

Linear Regression (Maulud & Abdulazeez, 2020): Linear regression is a algorithm that assumes a linear relationship between the independent variables and the dependent variable. It is suitable for modelling scenarios where there is a linear association between employee performance predictors and the target variable. Its simplicity and interpretability make it a commonly used baseline model for regression tasks.

Random Forest Regressor (Munir et al., 2023): Random Forest is an ensemble learning algorithm that combines multiple decision trees to make predictions. It is robust against overfitting and handles both categorical and numerical data well. Random Forest can capture complex interactions and non-linear relationships in the data, making it a suitable choice when the relationship between predictors and employee performance is non-linear.

XGBoost Regressor (Shahani et al., 2021): XGBoost (Extreme Gradient Boosting) is a gradient boosting algorithm known for its performance and accuracy. It combines weak prediction models (decision trees) sequentially, iteratively correcting the errors of the previous models. XGBoost handles missing values, feature interactions, and non-linear relationships effectively. Its ability to handle complex relationships and handle high-dimensional data makes it suitable for predicting employee working performance.

Decision Tree Regressor (Sishi & Telukdarie, 2021): Decision trees partition the data based on feature values and make predictions at each leaf node. They are interpretable and can capture non-linear relationships, interactions, and complex decision boundaries. Decision trees are useful when there are categorical predictors or when interpretability is a priority in employee performance prediction.

MLP Regressor (Dutt & Saadeh, 2022): MLP (Multi-Layer Perceptron) is a type of artificial neural network that consists of multiple layers of interconnected nodes. It can learn complex patterns and non-linear relationships in the data. MLP Regressor is suitable for capturing intricate interactions and hidden patterns in employee performance data.

SGD Regressor (Singh, 2022): SGD (Stochastic Gradient Descent) is an optimization algorithm used with various machine learning models, including linear regression and support vector machines. It is efficient and scalable, making it suitable for large datasets. SGD Regressor is chosen when there is a need for fast training and prediction in employee performance prediction.

Voting Regressor (Erdebilli, 2022): Voting Regressor combines multiple individual regression models by aggregating their predictions. It leverages the wisdom of the crowd, as the combined predictions tend to be more accurate and robust. Voting Regressor is suitable when there is uncertainty about which individual model would perform best in predicting employee working performance.

**Data Description Analytics and Visualization**

Descriptive analytics plays a crucial role in employee working performance prediction by providing insights into the characteristics, trends, and patterns of the data. In the context of employee working performance prediction, descriptive analytics helps to understand the distribution of attributes related to employee performance, such as working hours, productivity targets, and performance indicators. Descriptive analytics were also used to explore the relationships and associations between different variables. Correlation analysis helps identify the strength and direction of relationships between variables, providing insights into how factors like working hours or productivity targets may be related to employee performance outcomes.

To enhance the interpretation and presentation of the findings, Tableau software which is the data visualization tool has been utilized. Tableau enabled to create interactive and visually appealing visualizations effectively communicate the insights derived from the data. With Tableau's comprehensive range of visualization options, the charts, graphs, and interactive visual representations have been created that enhanced the understanding of the data and facilitated the exploration of complex relationships and patterns.

# EXPERIMENTAL RESULTS

This section presents analysis of the findings obtained from our research on employee working productivity prediction. This section encompasses the outcomes of three key components: Feature Ranking, Employee Working Productivity Prediction Model Performance, and Descriptive Analytics and Visualization

## Feature Ranking

This section presents the findings and analysis of a study that utilized various feature ranking algorithms, including InfoGainAttributeEval, GainRatioAttributeEval, CorrelationAttributeEval, OneRAttributeEval, and ReliefF. The purpose of feature ranking in employee performance prediction is to identify the key factors or attributes that have the most significant impact on predicting an employee's performance. By ranking the features, researchers and organizations can determine which variables are most relevant and influential in determining employee performance outcomes.

### *InfoGainAttributeEval*

![Chart, bar chart Description automatically generated](64c1f6c420a9d_media/media/image1.png)

**Figure** **1.** Feature Ranking using InfoGainAttributeEval

Based on the ranked attribute scores obtained using InfoGainAttributeEval shown in Figure 16, it can be summarized that the feature or attribute "Department" has the highest information gain score, indicating its strong influence on the target variable. "Actual Productive hours" and "Internet Speed" also have relatively high scores, suggesting their significant contribution to the prediction. Attributes like "Targeted Productive Hours," "COVID adoption month," and "WFH Challenges" show moderate influence. Attributes such as "Overtime hours," "Idle hours," "Daily Working Hours," and "Total Staff" have lower scores but still provide some predictive power. Other attributes have relatively lower information gain scores, indicating less influence on the target variable. Therefore, more attention should be given to the attributes with higher scores, such as "Department," "Actual Productive hours," and "Internet Speed," for accurate prediction. However, the context and specific requirements of the analysis should be considered while interpreting these attribute rankings.

###  ***CorrelationAtributeEval***

![Chart, bar chart Description automatically generated](64c1f6c420a9d_media/media/image2.png)

**Figure** **2.** Feature Ranking using CorrelationAtributeEval

Based on the attribute rankings obtained using CorrelationAttributeEval, it can be summarized that attributes such as "Actual Productive hours," "Internet Speed," and "COVID adoption month" exhibit the highest correlation scores, indicating a strong relationship with the target variable. Attributes like "Targeted Productive Hours," "Idle hours," "Overtime hours," and "Daily Working Hours" also demonstrate notable correlations. The attribute "Department" shows moderate correlation, while attributes such as "Working Preference," "Age," "Monthly Working Days," and "Total Staff" exhibit relatively lower correlations. Other attributes, including "WFH Challenges," "Internet Problem Solution," "Working Promotion," and "ï»¿Gender," demonstrate weaker correlations. These rankings provide insights into the potential influence of each attribute on the target variable, but the specific context and requirements of the analysis should be considered for a comprehensive interpretation.

###  ***GainRatioAttributeEval*** 

###  

![Chart, timeline, bar chart Description automatically generated](64c1f6c420a9d_media/media/image3.png)

**Figure** **3.** Feature Ranking GainRatioAttributeEval

The feature ranking using the GainRatioAttributeEval algorithm illustrated in Figure 18 provides insights into the importance of different attributes for predicting employee performance. The top-ranked attribute is Actual Productive Hours (0.2655), followed closely by Internet Speed (0.2495) and Targeted Productive Hours (0.2476). These attributes play a crucial role in determining employee performance outcomes. Other significant attributes include COVID Adoption Month (0.2204) and Department (0.1538). Lower-ranked attributes such as Idle Hours (0.0836), Overtime Hours (0.0773), Total Staff (0.066), WFH Challenges (0.0635), Daily Working Hours (0.0624), Age (0.0482), Monthly Working Days (0.0475), and Internet Problem Solution (0.0462) also contribute to performance prediction, albeit to a lesser extent. These rankings highlight the key attributes that strongly influence employee performance, guiding organizations in improving productivity and optimizing their workforce.

***OneRAttributeEval***

![Chart, bar chart Description automatically generated](64c1f6c420a9d_media/media/image4.png)

**Figure** **4.** Feature Ranking using OneRAttributeEval

The feature ranking using the OneRAttributeEval algorithm shown in Figure 19 provides valuable insights into the attributes' importance for predicting employee performance. The top-ranked attribute is Actual Productive Hours (60.194), indicating its strong influence on employee performance outcomes. Following closely is Internet Speed (56.311), emphasizing the significance of a fast and reliable internet connection. Daily Working Hours (49.515) also holds importance, reflecting the impact of the duration of work on performance. Additionally, Targeted Productive Hours, COVID adoption month, and Department (all with 48.544) are identified as equally influential factors. Monthly Working Days (47.573) further contributes to performance prediction. Other attributes, such as Working Promotion, Gender, WFH Challenges, Working Preference, Internet Problem Solution, Total Staff, Idle Hours, Overtime Hours, and Age, have a relatively lesser impact. These findings provide organizations with a deeper understanding of the attributes that strongly correlate with employee performance, enabling them to focus on optimizing these factors to enhance productivity.

***ReliefF***

![Chart, timeline, bar chart Description automatically generated](64c1f6c420a9d_media/media/image5.png)

**Figure** **5**. Feature Ranking using ReliefF

The feature ranking using the ReliefFAttributeEval algorithm provides insights into the attributes' importance for predicting employee performance. The ranked attributes, along with their corresponding ReliefF scores, indicate their impact on performance prediction. The top-ranked attribute is Actual Productive Hours (0.22651), highlighting its strong influence on employee performance outcomes. Targeted Productive Hours (0.18474) and Internet Speed (0.17843) also demonstrate significant importance in predicting performance. COVID adoption month (0.13916) follows closely, indicating the influence of the time when employees adapted to pandemic-related changes. Conversely, attributes such as Age (-0.00402), Daily Working Hours (-0.00486), Monthly Working Days (-0.00692), and Gender (-0.01668) show negative scores, suggesting a weaker correlation with performance. Other attributes such as Idle Hours, Total Staff, Working Promotion, Working Preference, Department, Overtime Hours, WFH Challenges, Internet Problem Solution, and Department contribute to performance prediction to varying degrees. These findings enable organizations to focus on the most influential attributes when developing strategies to enhance employee performance and productivity.

Based on the findings from various feature ranking methods, certain attributes stand out as highly influential for predicting employee performance. According to InfoGainAttributeEval, "Department" exhibits the highest information gain score, along with "Actual Productive hours" and "Internet Speed" showing significant contributions. CorrelationAttributeEval highlights strong relationships between performance and attributes like "Actual Productive hours," "Internet Speed," and "COVID adoption month." GainRatioAttributeEval identifies "Actual Productive Hours," "Internet Speed," and "Targeted Productive Hours" as the top attributes, while ReliefFAttributeEval ranks "Actual Productive Hours," "Internet Speed," and "Targeted Productive Hours" as most important. Other attributes also contribute to performance prediction to varying extents. In conclusion, attributes such as "Department," "Actual Productive hours," and "Internet Speed" should receive special attention when formulating strategies to enhance employee performance, considering the context and requirements of the analysis for comprehensive interpretation.

**Employee Working Productivity Prediction Model Performance**

The experiment aimed to predict employee working performance using various machine learning regression algorithms: Linear Regression (LR), Random Forest Regressor (RF), XGBoost Regressor (XB), Decision Tree Regressor (TREE), MLP Regressor (MLP), SGD Regressor (SGD), and Voting Regressor (KNN). The performance of each model was evaluated using Mean Absolute Error (MAE) and Mean Squared Error (MSE) as shown in Figure 21.

![](64c1f6c420a9d_media/media/image6.png)

**Figure** **6.** Employee Working Performance Prediction Results

As shown in Figure 21, Linear Regression attempts to fit a linear relationship between the features and the target variable. The achieved MAE of 0.4878 suggests that, on average, the predicted values deviate from the actual values by approximately 0.49 units. The MSE of 0.4682 indicates the mean squared difference between the predicted and actual values, with higher weight given to larger errors. Next, Random Forest Regressor is an ensemble learning method that builds multiple decision trees and averages their predictions. The achieved MAE of 0.5406 indicates that the model's predictions deviate from the actual values by around 0.54 units on average. The MSE of 0.4631 suggests the mean squared difference between predicted and actual values. XGBoost is a gradient boosting algorithm known for its performance in various machine learning tasks. The achieved MAE of 0.5438 shows that the model's predictions have an average deviation of approximately 0.54 units from the actual values. The MSE of 0.5052 indicates the mean squared difference between predicted and actual values. MLP Regressor is a neural network-based model. The achieved MAE of 0.7519 suggests relatively high prediction errors. The MSE of 1.1868 indicates larger mean squared differences between predicted and actual values. Voting Regressor combines the predictions of multiple models. The achieved MAE of 0.5018 indicates a reasonably low average deviation between predicted and actual values. The MSE of 0.4559 suggests a relatively smaller mean squared difference.

The experiment's findings on predicting employee working performance using different regression algorithms revealed that the achieved Mean Absolute Error (MAE) values ranged from approximately 0.488 to 0.752, indicating the average deviation between predicted and actual values. The Mean Squared Error (MSE) values ranged from around 0.451 to 1.187, representing the mean squared difference between predicted and actual values. Despite the variation in MAE and MSE, the overall performance of all regression models was suboptimal, suggesting that the selected algorithms, along with the feature set, may not be the most suitable for accurately predicting employee performance. This highlights the complexity of the task and the need for alternative modeling approaches, feature engineering, data quality improvements, domain knowledge incorporation, and potentially more advanced machine learning techniques to achieve more accurate predictions for employee performance in this specific context.

**Descriptive Analytics and Visualization**

In this section, we present the visualization component of our research. The analysis in this section focuses on exploring various key aspects related to employee working productivity and its interactions with different factors within the organization. The outlined subsections include Age Versus Productivity, Department versus Actual Productivity Hours, Department versus Working Productivity during the COVID-19 pandemic, Internet Connections Performance versus Working Productivity, Age versus Idle Hours, Age versus Overtime, Age versus Working Promotion, and Age versus Work From Home (WFH) Challenges.

***Age Versus Productivity***

![](64c1f6c420a9d_media/media/image7.png)

The data provides insights into the average productive hours before and after the pandemic for different age groups. The age groups considered in the analysis are 25 to 30, 36 to 45, 46 to 55, and more than 55.

Before the pandemic, the average productive hours varied among the age groups. The age group of 25 to 30 exhibited an average of 3.4412 productive hours. The 36 to 45 age group had slightly lower average productive hours at 3.3333. For the 46 to 55 age group, the average productive hours increased to 3.5714. The age group of more than 55 also showed a relatively high average of 3.5 productive hours.

After the pandemic, there was a slight overall increase in average productive hours for all age groups. The 25 to 30 age group experienced an average of 3.6176 productive hours, reflecting a moderate increase from the pre-pandemic period. Similarly, the 36 to 45 age group saw an increase to 3.5 average productive hours. The 46 to 55 age group exhibited a further increase to 3.7429 productive hours. The age group of more than 55 had the highest average productive hours after the pandemic, with an average of 3.8.

The findings suggest that, in general, there was a slight increase in average productive hours across age groups following the pandemic. However, it is important to consider other factors, such as the specific industry or work circumstances, which could influence these changes in productivity. Further analysis and examination of additional variables are necessary to gain a comprehensive understanding of the impact of the pandemic on productivity within different age groups.

***Department versus Actual Productivity Hours***

![Chart, bar chart Description automatically generated](64c1f6c420a9d_media/media/image11.png)

Figure : Employee Productivity by Department before Pandemic Covid-19

The provided data presents the average productive hours before COVID-19 across different departments within an organization. The analysis reveals variations in productivity levels observed among these departments during that period.

Unit Psikologi & Kaunseling and Unit Peperiksaan & Penilaian demonstrate the highest average productive hours, both recording a value of 4. These departments indicate a relatively high level of productivity before COVID-19.

Jabatan Pendidikan Islam & Moral, Unit Praktikum, and Pengurusan Tertinggi also display strong average productive hours, with values of 3.777777778, 3.611111111, and 4, respectively. These departments showcase above-average productivity levels during the pre-COVID-19 period.

On the other hand, Unit Kokurikulum, Jabatan Hal Ehwal Pelajar, Unit Khidmat Pengurusan, and Pusat Sumber exhibit lower average productive hours, with values ranging from 2.888888889 to 3.206349206. These departments suggest a relatively lower level of productivity compared to others.

The variations in average productive hours across departments highlight potential differences in workloads, effectiveness, or other factors influencing productivity. These findings provide insights into departmental performance and can guide resource allocation and performance improvement initiatives within the organization.

Organizational leaders can utilize this data to identify departments with high productivity levels as benchmarks and explore strategies to enhance productivity in departments with lower averages. By leveraging these insights, organizations can optimize productivity, allocate resources effectively, and foster a culture of continuous improvement in the workplace.

***Department versus Working Productivity during COVID-19 pandemic***

![Chart, bar chart Description automatically generated](64c1f6c420a9d_media/media/image12.png)

Figure : Employee Productivity by Department after Pandemic Covid-19

The provided data compares the average productive hours during the COVID-19 period with the previous data before the pandemic. It reveals the changes in productivity levels across different departments within the organization.

During COVID-19, several departments experienced a decrease in average productive hours compared to the pre-pandemic period. Unit Asrama, Unit Peperiksaan & Penilaian, Jabatan Pendidikan Islam & Moral, Jabatan Sains Sosial & Teknologi, Jabatan Pendidikan Jasmani & Kesihatan, Jabatan Hal Ehwal Pelajar, Jabatan Bahasa, and Unit Teknologi Maklumat & Komunikasi all maintained the same average productive hours as before.

However, some departments demonstrated an increase in average productive hours during COVID-19. Notably, Unit Kokurikulum, Jabatan Sains, Jabatan Matematik, Jabatan Pengajian Melayu, Jabatan Kecemerlangan Akademik, Pusat Sumber, Unit Praktikum, Unit Psikologi & Kaunseling, Unit Pengurusan Kewangan & Akaun, Pengurusan Tertinggi, Unit Khidmat Pengurusan, and Unit Pembangunan Latihan experienced a rise in productivity.

These findings suggest that certain departments were able to maintain or even improve their productivity levels during the challenging circumstances of the COVID-19 pandemic. The data highlights the resilience and adaptability of these departments in navigating the changes and optimizing their productivity. It also indicates areas where additional support or strategies may be needed to enhance productivity for departments that experienced a decline in average productive hours during the pandemic.

***Internet Connections Performance versus Working Productivity***

![Chart, bar chart Description automatically generated](64c1f6c420a9d_media/media/image13.png)

Figure : Internet Connectivity Performance versus Employee Productivity

The findings from the employee performance prediction data shed light on the critical relationship between internet connectivity performance and average productivity hours. The analysis reveals distinct variations in productivity levels based on the quality of internet connectivity. Employees with an average internet connectivity performance displayed an average productivity of 3.333 hours. This suggests that while they may meet the baseline requirements, their productivity remains average and leaves room for improvement. On the other hand, employees with good internet connectivity exhibited a higher average productivity of 3.738 hours, indicating a positive correlation between better internet performance and increased productivity.

Conversely, employees facing poor internet connectivity experienced a decline in average productivity, with an average of only 3.167 hours. This raises concerns about the impact of unreliable or slow internet connections on employee performance and overall output. The limited connectivity likely hampers their ability to efficiently complete tasks, resulting in suboptimal productivity levels.

Of particular interest is the significant disparity observed in employees with very good and very poor internet connectivity. Those with very good internet connectivity demonstrated an impressive average productivity of 4.643 hours, showcasing the positive influence of a robust and high-speed connection on work output. Conversely, employees dealing with very poor internet connectivity struggled significantly, with an alarming average productivity of only 2.5 hours. Such findings underscore the detrimental effects of severely limited or unreliable internet connections on employee productivity.

These results emphasize the critical role of internet connectivity in shaping employee performance. Organizations should recognize the importance of providing employees with stable and efficient internet connections to optimize productivity and ensure seamless workflow. Addressing issues related to internet connectivity, such as bandwidth limitations or network stability, becomes crucial for fostering a productive work environment.

Investing in infrastructure improvements, exploring alternative connectivity solutions, or implementing measures to mitigate internet-related challenges can lead to significant productivity gains. By prioritizing and enhancing internet connectivity, organizations can empower employees to perform at their best, ultimately driving overall business success.

***Age versus Idle Hours***

![Chart, bar chart Description automatically generated](64c1f6c420a9d_media/media/image14.png)

Figure : Age versus Idle hours

The analysis focuses on the relationship between age and idle hours among employees, shedding light on potential trends and patterns. The data is categorized into four age groups: 25 to 30, 36 to 45, 46 to 55, and more than 55. Employees in the age group of 25 to 30 exhibited an average of 1.7941 idle hours, indicating a relatively lower level of unproductive time. However, it is crucial to interpret this finding with caution, as idle hours can be influenced by various factors, including job roles, tasks, and work environments. A deeper examination of these contextual elements is necessary to better understand the implications of age on idle hours.

The 36 to 45 age group showed slightly higher idle hours, with an average of 2.0417. This finding suggests that individuals within this age range may experience slightly more unproductive time during work. The factors contributing to these idle hours could range from distractions, work interruptions, or potential inefficiencies in task management. Organizations should consider exploring potential strategies to address these factors and optimize productivity.

Similarly, the 46 to 55 age group exhibited a comparable average of 2.0571 idle hours. The similarity in idle hours between the 36 to 45 and 46 to 55 age groups may indicate consistent challenges in managing work-related tasks and minimizing unproductive time. Organizations should focus on identifying underlying causes for these idle hours and implement targeted interventions to boost productivity.

Notably, the age group of more than 55 demonstrated a slightly lower average of 1.7 idle hours. This finding is intriguing and suggests that individuals in this age group may have developed effective time management strategies or possess a higher level of task efficiency. However, it is essential to consider that these observations are based solely on idle hours and may not capture the full productivity dynamics within the workplace.

To gain a comprehensive understanding of the relationship between age and idle hours, it is crucial to consider additional factors such as job responsibilities, workloads, work environment, and individual work habits. Conducting more detailed analyses and considering these contextual factors will facilitate a more thorough and critical assessment of the impact of age on idle hours and overall productivity levels.

***Age versus Overtime***

![](64c1f6c420a9d_media/media/image15.png)

Figure : Age versus Overtime

The data provided presents the distribution of total overtime hours based on different age groups. Among the age groups, individuals between 46 to 55 years accounted for the largest percentage of total overtime hours, representing 36.18% of the overall workload. The age group of 25 to 30 years contributed significantly to the total overtime hours, making up 30.65%. Meanwhile, individuals aged 36 to 45 years accounted for 24.62% of the total overtime hours. Those above the age of 55 contributed the least, comprising only 8.54% of the overall overtime workload. These findings suggest that individuals in the age range of 46 to 55 years are more likely to work longer hours and contribute a substantial portion of the total overtime workload. Conversely, individuals above the age of 55 have a smaller presence in overtime work.

***Age versus Working Promotion***

![](64c1f6c420a9d_media/media/image16.png)

Figure : Age versus Working Promotion

The analysis examines the relationship between age and promotion rates, specifically focusing on the number of individuals who have been promoted versus those who have never been promoted. The data reveals varying promotion rates across different age groups, shedding light on potential patterns and trends.

In the age group of 25 to 30, approximately 47.06% (16 out of 34) of individuals have been promoted, while approximately 52.94% (18 out of 34) have never been promoted. This relatively balanced distribution suggests that individuals in this age bracket have had comparable opportunities for career advancement. However, it is essential to consider other factors such as performance, qualifications, and experience, as they may contribute to the observed promotion rates.

For the age group of 36 to 45, approximately 33.33% (8 out of 24) have been promoted, while approximately 66.67% (16 out of 24) have never been promoted. The relatively low promotion rate within this age range raises questions about potential barriers or limitations to career progression. It is worth investigating whether factors such as competition, organizational structure, or career stagnation contribute to the lower promotion rates observed in this group.

Similarly, in the age group of 46 to 55, approximately 34.29% (12 out of 35) have been promoted, while approximately 65.71% (23 out of 35) have never been promoted. This finding suggests a similar trend to the 36 to 45 age group, indicating potential challenges or limitations to career advancement within this age range. Organizations should explore possible factors such as skill gaps, biases, or limited growth opportunities that may contribute to the observed promotion rates.

Interestingly, in the age group of more than 55, approximately 30% (3 out of 10) have been promoted, while approximately 70% (7 out of 10) have never been promoted. This finding raises concerns about potential age-related biases or limited promotional opportunities for experienced professionals. Further investigation is necessary to identify underlying factors that contribute to this disparity and to ensure fair and equal career advancement opportunities for employees in this age group.

It is important to note that promotion decisions are influenced by various factors beyond age, such as performance, qualifications, and organizational policies. A comprehensive analysis considering these factors is required to gain a deeper understanding of the relationship between age and promotion rates. This would enable organizations to identify and address any potential biases or barriers to career progression, fostering a more inclusive and equitable work environment.

***Age versus Work From Home (WFH) Challenges***

![Chart, bar chart Description automatically generated](64c1f6c420a9d_media/media/image17.png)

Figure : Age versus WFH Challenges

Based in the graph shown in Figure 29, among individuals aged 25 to 30 working from home, common challenges include family member disruptions (11.65%) and internet connection problems (8.74%). Additionally, some individuals struggle with maintaining focus (2.91%) and understanding assigned tasks (1.94%). A smaller portion faces challenges related to the workspace environment (non-conducive place) and other unspecified difficulties. For individuals aged 36 to 45 working from home, maintaining focus (6.80%) and dealing with family member disruptions (5.83%) are the primary challenges. Some individuals also face difficulties understanding assigned tasks (2.91%) and have issues related to their workspace environment (non-conducive place). A smaller portion encounters other unspecified challenges. Among individuals aged 46 to 55 working from home, challenges include family member disruptions (8.74%) and internet connection problems (6.80%). Maintaining focus (4.85%) and understanding assigned tasks (2.91%) are also reported challenges. Some individuals face difficulties with their workspace environment (non-conducive place), and others encounter other unspecified challenges. For individuals aged more than 55 working from home, family member disruptions (0.97%) and internet connection problems (4.85%) are the primary challenges. Some individuals also encounter difficulties with their workspace environment (non-conducive place) and have other unspecified challenges.

In conclusion, the insights from the breakdown of WFH challenges across different age groups provide valuable guidance for employers to create targeted and effective solutions. By understanding the specific challenges faced by employees in different age brackets, organizations can implement measures to enhance productivity, work-life balance, and employee well-being in the context of remote work. This, in turn, can foster a more positive and productive work environment, ultimately benefiting both the employees and the organization as a whole.

**CONCLUSIONS AND RECOMMENDATIONS**

This research has developed a predictive model using data mining techniques to determine employee productivity within organizations by using various machine learning prediction models and identified the important attributes affecting employee working performance based on several machine learning ranker algorithms. Descriptive analytics techniques aid in visualizing, plotting, and analysing the data, extracting valuable insights and understanding attributes correlations. Among the evaluated models, the linear regression model emerges as the most accurate predictor for employee productivity in the given organizational context, with MAE and MSE values of 0.4878 and 0.4682, respectively. In light of the research findings, it is recommended that organizations consider adopting linear regression for predicting employee productivity. Additionally, implementing effective data visualization methods will help gain deeper insights from the available data. Further research can focus on exploring more advanced machine learning algorithms, incorporating time-series analysis for temporal dependencies, and expanding data collection from diverse organizational settings to improve the generalizability of predictive models.

Subsequently, based on the findings from various attributes ranking algorithm, it is evident that certain attributes play a crucial role in predicting employee performance. Notably, "Department," "Actual Productive hours," and "Internet Speed" emerged as highly influential factors across multiple ranking techniques. These attributes consistently displayed strong correlations with employee performance, indicating their significance in determining productivity levels. "Department" was identified as the attribute with the highest information gain, suggesting that the department in which an employee works has a substantial impact on their performance. Understanding the variations in performance across different departments can aid in targeted interventions and resource allocation. "Actual Productive hours" consistently appeared as a key attribute in all ranking methods. This highlights the importance of employees' actual productive hours in assessing their performance accurately. Organization should focus on optimizing the time employees spend on productive tasks to enhance overall productivity. The attribute "Internet Speed" also featured prominently in the rankings, signifying the relevance of a stable and efficient internet connection in facilitating employee performance. Organizations should invest in robust internet infrastructure and support systems to ensure seamless remote work and avoid productivity hindrances. Additionally, the "COVID adoption month" attribute showed significant relevance, implying that the duration of adaptation to pandemic norms impacted employee performance. Organizations should consider the challenges faced during the pandemic and implement measures to support employees during such transitions. While other attributes contributed to performance prediction to varying degrees, these top-ranking attributes warrant special attention when devising strategies to improve employee performance.

Finally, the data visualization has provided valuable insights into various aspects of employee performance, including average productive hours before and after the pandemic for different age groups, departmental performance, internet connectivity's impact on productivity, age-related trends in idle hours, distribution of total overtime hours by age groups, and promotion rates across different age brackets. The findings indicate changes in productivity levels and shed light on factors influencing employee performance. The data suggests a slight increase in average productive hours across all age groups after the pandemic, but further analysis is needed to understand the full impact of the pandemic on productivity, considering industry-specific circumstances. Variations in departmental performance highlight the need for targeted strategies to enhance productivity in specific departments. The relationship between internet connectivity and productivity emphasizes the importance of stable connections in optimizing employee performance. While age may have a slight impact on idle hours, contextual factors must be considered for a comprehensive understanding. Workforce planning can be informed by recognizing the propensity of the 46 to 55 age group to contribute the most to total overtime hours. Organizations should address age-related challenges and improve internet connectivity to foster an inclusive work environment and encourage work-life balance. Implementing fair promotion processes will ensure equitable opportunities for career advancement. By regularly assessing productivity data and implementing recommended strategies, organizations can enhance productivity, create a conducive work environment, and support employee well-being and growth.

**REFERENCES**

Alhaj, T. A., Siraj, M., Zainal, A., Elshoush, H. T., & Elhaj, F. (2016). *Feature Selection Using Information Gain for Improved Structural-Based Alert Correlation*. 1–18. https://doi.org/10.1371/journal.pone.0166017Atatsi, E. A., Stoffers, J., & Kil, A. (2019). Factors affecting employee performance: a systematic literature review. *Journal of Advances in Management Research*, *16*(3), 329–351. https://doi.org/10.1108/JAMR-06-2018-0052Billson, R., Schiel, A., Yu-bo, Z., & Ming-duo, Y. (2021). *Attributes selection using machine learning for analysing students ’ dropping out of university : a case study Attributes selection using machine learning for analysing students ’ dropping out of university : a case study*. https://doi.org/10.1088/1757-899X/1031/1/012055Dutt, M. I., & Saadeh, W. (2022). A Multilayer Perceptron ( MLP ) Regressor Network for Monitoring the Depth of Anesthesia. *2022 20th IEEE Interregional NEWCAS Conference (NEWCAS)*, 251–255. https://doi.org/10.1109/NEWCAS52662.2022.9842242Erdebilli, B. (2022). *Ensemble Voting Regression Based on Machine Learning for Predicting Medical Waste : A Case from Turkey*. 7–9.Hu, B. (2021). The application of machine learning in predicting absenteeism at work. *Proceedings - 2021 2nd International Conference on Computing and Data Science, CDS 2021*, 270–276. https://doi.org/10.1109/CDS52072.2021.00054Jayadi, R., Jayadi, R., Firmantyo, H. M., Dzaka, M. T. J., Suaidy, M. F., & Putra, A. M. (2019). *of Advanced Trends in Computer Science November and Employee Performance Prediction using Naïve Bayes*. *8*(6), 8–12.Li, M. G. T., Lazo, M., Balan, A. K., & De Goma, J. (2021). Employee performance prediction using different supervised classifiers. *Proceedings of the International Conference on Industrial Engineering and Operations Management*, 6870–6876.Maulud, D. H., & Abdulazeez, A. M. (2020). *A Review on Linear Regression Comprehensive in Machine Learning*. *01*(04), 140–147. https://doi.org/10.38094/jastt1457Monisaa Tharani, S. K., & Vivek Raj, S. N. (2020). Predicting employee turnover intention in ITITeS industry using machine learning algorithms. *Proceedings of the 4th International Conference on IoT in Social, Mobile, Analytics and Cloud, ISMAC 2020*, 508–513. https://doi.org/10.1109/I-SMAC49090.2020.9243552Munir, S., Seminar, K. B., Sukoco, H., & Buono, A. (2023). *The Use of Random Forest Regression for Estimating Leaf Nitrogen Content of Oil Palm Based on Sentinel 1-A Imagery*.Naz, K., Siddiqui, I. F., Koo, J., Khan, M. A., & Qureshi, N. M. F. (2022). Predictive Modeling of Employee Churn Analysis for IoT-Enabled Software Industry. *Applied Sciences (Switzerland)*, *12*(20). https://doi.org/10.3390/app122010495Patil, M. D. (2014). *Effective Classification after Dimension Reduction : A*. *4*(7), 1–4.Sabuj, H. H., Nuha, N. S., Gomes, P. R., Lameesa, A., & Alam, M. A. (2023). *Interpretable Garment Workers’ Productivity Prediction in Bangladesh Using Machine Learning Algorithms and Explainable AI*. 236–241. https://doi.org/10.1109/iccit57492.2022.10054863Saputra, W., & Purwitasari, D. (2022). Fatigue Management: Machine Learning Application for Predicting Mining Worker Fatigue. *2022 International Conference on Information Technology Research and Innovation, ICITRI 2022*, 117–122. https://doi.org/10.1109/ICITRI56423.2022.9970203Sarker, A., Shamim, S. M., Shahiduz, M., Rahman, Z. M., Shahiduz Zama, M., & Rahman, M. (2018). Employee’s Performance Analysis and Prediction using K-Means Clustering & Decision Tree Algorithm. *International Research Journal Software & Data Engineering Global Journal of Computer Science and Technology*, *18*(1), 7. https://computerresearch.org/index.php/computer/article/view/1660/1644Shahani, N. M., Zheng, X., Liu, C., & Hassan, F. U. (2021). *Developing an XGBoost Regression Model for Predicting Young ’ s Modulus of Intact Sedimentary Rocks for the Stability of Surface and Subsurface Structures*. *9*(October), 1–13. https://doi.org/10.3389/feart.2021.761990Singh, G. (2022). *Machine Learning Models in Stock Market Prediction*. *3075*(3), 18–28. https://doi.org/10.35940/ijitee.C9733.0111322Sishi, M., & Telukdarie, A. (2021). *The Application of Decision Tree Regression to Optimize Business Processes*. *Dm*, 48–57.Tambde, A., & Motwani, D. (2019). Employee churn rate prediction and performance using machine learning. *International Journal of Recent Technology and Engineering*, *8*(2 Special Issue 11), 824–826. https://doi.org/10.35940/ijrte.B1134.0982S1119Zainudin, M. N. S., Sulaiman, N., Mustapha, N., Perumal, T., & Mohamed, R. (2018). Two-stage feature selection using ranking self-adaptive differential evolution algorithm for recognition of acceleration activity. *Turkish Journal of Electrical Engineering and Computer Sciences*, *26*(3), 1378–1389. https://doi.org/10.3906/elk-1709-138
