**Prediction of depression using Machine learning tools taking consideration of oversampling.**

\(\text{Md.\ Murad\ Hossain}^{1,2}\), \({}^{2},\) \(\text{Md.\ Amzad\ Hossain\ }^{3}\) and\(\ \text{\ Muhammad\ Saad\ Amin\ }^{4}\)

<table>
<tbody>
<tr class="odd">
<td><table>
<tbody>
<tr class="odd">
<td><h3 id="section"></h3></td>
</tr>
</tbody>
</table></td>
</tr>
</tbody>
</table>

1.  Modeling and Data science, University of Turin, Via Verdi,8-10124 Turin, Italy.

2.  Department of Statistics, Bangabandhu Sheikh Mujibur Rahman Science and Technology University, Gopalganj 8100, Bangladesh.

3.  Department of Information and Communication Engineering, Noakhali Science and Technology University, Noakhali 3814, Bangladesh.

4.  Department of Computer Science, University of Turin, Via Verdi,8-10124 Turin, Italy.

**Abstract**

Depression is a psychiatric condition characterized by a persistent sense of sadness and dullness. It is also known as a severe burdensome problem or clinical sorrow, and it impacts how a person feels, thinks, and behaves and triggers a slew of emotional and physical issues. Depression is spreading at an unprecedented pace across the world. Various components are liable for this issue, and many related sicknesses are expanding because of this infection. Gloom is not just at risk for well-being perils, yet also produces perilous social offense, like self-destruction and family misuse. In this study, We used machine learning methods such as Random Forest (RF), Logistic Regression (LR), and Naive Bayes (NB). We also used accuracy, precision, recall, and F1-score to survey the exhibition assessment of arrangement results. These machine learning algorithms developed and analyzed Confusion matrices through data augmentation to assess the classification performance. This study used machine learning technologies to predict depression and revealed the significance of the trait. Then we have tried to utilize an oversampling technique that shown the distinction in model execution. Indeed, we wanted to see how well the recommended machine learning algorithms performed before and after rebalancing standardized data. In my suggested framework, the RF classifier performed better with 89% accuracy and 90% precision than other models.

**Keywords: Depression, Machine Learning, Classification, Accuracy, oversampling.**

**Introduction**

People are, naturally, turning out to be goal-oriented these days and look for each reasonable chance to develop professionally. Nervousness, sadness, stress, disappointment, and dissatisfaction have become so ordinary that individuals presently trust them to be vital in daily life. Sau et al. \[1\] explains that depression is a common psychiatric disease. According to the World Health Organization, it afflicted more than 300 million individuals globally. The seriousness of the epidemic has caused many health professionals to concentrate their studies on it. Choi et al. \[2\] use six predictive models and two missing value imputation methods for constructing an inner depression foresight model for elderly individuals. Several machine-learning techniques, including logistic regression, a ridge estimator, random forest, and even two faulty data inference approaches, were utilized to examine the possibility of using machine-learning methods for a widely available dataset. But they have not considered any tuning method for enhancing the performance of the model. Kipli et al. \[3\] the entire exploration sees machine learning algorithms for predicting depression by systematically defining relevant data properties uses three extraction approaches and one composite methodology, along with stat tech analysis tools, selects subsets of characteristics to eliminate redundant attributes. Dipnall, Joanna F. et al. \[4\] to find key biomarkers linked to depression they used a three-step methodology that included several accusations, an LR model with machine learning, and an enhanced regression model. They examined machine learning techniques that automatically select significant data qualities for stress detection using these algorithms. Hasanzadeh et al. \[5\] look at some of the most basic classification algorithms for brain scan diagnosis and detection. Ultimately, problems, prospective paths, and possible drawbacks address relevant to major depression disorder biomarker detection. Jiménez-Serrano et al. \[6\] aim to create classification techniques for identifying the risk of postpartum depression within a week after delivery, allowing early detection, and creating a mobile health app for the Android platform. They focused on the best model for both young moms and physicians who want to keep track of their patient's test results. Priya et al. \[7\] utilize Machine learning algorithms to conjecture tension, gloom, stress, and severity levels of anxiety. Information acquired from working ruined people from different foundations, and five separate ml calculations extended their appearance on five degrees of solidarity. Zhang, Yiye, et al. \[8\] proposed a generic application for forecasting postpartum depression risk based on data from e-health reports. A cycle of data extraction, unravelling, and information science use to choose a modest number from the electric security dataset to submitted unwavering quality and future place of time hazard expectations. Xie, Zidian, et al. \[9\] analyze the cross-sectional findings using the 2014 Cognitive Risk Factor Monitoring Method, uses the dataset to predict type 2 diabetes through the SVM classification scheme, DT, LR, RF, neural network (NN), and Gaussian Naïve Bayes optimization techniques. They also used univariate analysis and multivariate calibrated logistic regression analysis to evaluate the connections of adverse outcomes with type 2 diabetes. Srividya et al. \[10\] use SVM, DT, NB classification techniques, and K-nearest neighbors are among the information retrieval protocols mentioned. The authors look at the effects of the above machine learning algorithms on groups and recommend future research. Iliou, Theodoros, et al. \[11\] identify a feature extraction approach. They created a pre-processing data system for detecting depression types. Also used Several statistical and machine learning approaches test through the ten-fold validation set framework. Kumar et al. \[12\] evaluated significant zones in the dataset that majorly affected ordering a patient's psychological steadiness. They applied correlation analysis and selected machine learning-based strategies. Hosseinifard et al. \[13\] demonstrated how to use non - linearity Electroencephalogram (EEG) signal analysis to distinguish between depression patients and control. The models used to characterize the groups are KNN, discriminant classification, and logistic regression. Islam, Md Rafiqul, et al. \[14\] aimed to conduct a mental illness analysis on Facebook data obtained from a community internet site. The author proposes machine learning as an effective and helpful method to analyze the influence of emotion classification. Khalil et al. \[15\] evaluate training machine learning strategies for detecting stress behaviors with data modifications, state-of-the-art ensemble learning. The authors also attempted to adopt a compact approach for class label assignment. Farima et al. \[16\] look at how content models can forecast postpartum depression using knowledge from social media sites. Powerful machine learning approaches derive linguistic features from consumer textual social media posts, categorize them as ordinary, traumatic, and simulate a postnatal depression model. Ghandeharioun, Asma, et al. \[17\] the utility of the data analysis model for predicting the Hamilton Rating Scale develops and tests using observational data. They explored how they produced features and converted them and how they imputed missed therapeutic scores from self-reported assessments and estimated depression intensity from multiple persistent sensors. Nouretdinov, Ilia, et al. \[18\] proposed a new deterministic classification method for generating measures of confidence based on positron emission tomography results. They also described the application of transductive conformal predictor (TCP) to magnetic resonance imaging (MRI) images.  Bhakta et al. \[19\] choices three tests to conduct and comparing five machine learning classifiers. After comparing many approaches, they suggested the best method for predicting depression in older adults. However, they have considered only adult people in their study. Khodayari-Rostamabad, Ahmad, et al. \[20\] tried to find a computational algorithms strategy for predicting exposure to diagnosis with a serotonin receptor medication in people with psychosis analysis. Also, a small proportion of the most markers select from a vast number of eligible items. Koutsouleris, Nikolaos, et al. \[21\] attempted to show whether public and status behaviour determinants can establish in people with emotionally elevated conditions for disease using diagnostic, tomography, and mixed data science.

Nonetheless, a few the publications listed above use measurable research to characterize wretchedness, such as fundamental structures, dispersions, and regression models. A handful of publications describe the bio-maker underlying anxiety by studying recent comments. Indeed, employing various validation and authenticating procedures, specific assessment papers portrayed stress. Several articles have described how covid-19 enhanced depression during society's lockdown time. Nonetheless, a small number of articles incorporated machine learning techniques in the oversampling model. There is no such examination accessible that has endeavoured to distinguish and estimate pressure utilizing machine learning software. To forecast depression, we used contemporary machine learning techniques and the crucial element that causes depression and suggest an oversampling strategy throughout this study.

Additionally, we endeavoured to show a better performance using oversampling methods to the current machine learning tools. Throughout this essay, we have evaluated the utility of various analysis tools from the perspectives of description, validation, and accuracy. Most of the authors address the factors related to depression, and few authors finding some impact of depression. However, no one explains how to handle imbalanced data to predict target variables using a machine learning algorithm. In this study, we tried to run imbalanced data to predict depression and improve model accuracy.

**MATERIALS AND METHODS**

**Data Description**

The data set we have used for our project report, collected from Kaggle. The link to the data is https://www.kaggle.com/diegobabativa/depression. There are a few different variables in this dataset such as sex, age, Married, Number\_children, education\_level, total\_members, gained\_asset, durable\_asset, save\_asset, living\_expenses, other\_expenses, incoming\_salary, incoming\_own\_farm, incoming\_business, incoming\_no\_business, incoming\_agricultural, farm\_expenses, labor\_primary, lasting\_investment, no\_lasting\_investmen, target.We utilized the R package version 4.03 for data management and analysis.

**Table 1:** Arrangement of inquiries in the research project on depression.

| **S/N** | **Feature**            | **Description**                                                        |
| ------- | ---------------------- | ---------------------------------------------------------------------- |
| 1       | sex                    | Gender of the respondent                                               |
| 2       | Age                    | The age of the respondent                                              |
| 3       | Married                | The respondent's Marital Status                                        |
| 4       | Number\_children       | Number of kids of the Participant's                                    |
| 5       | education\_level       | Educational attainment                                                 |
| 6       | total\_members         | Total family members of the Respondent                                 |
| 7       | gained\_asset          | Acquired assets of the Respondent                                      |
| 8       | durable\_asset         | Sustainable resource of the Respondent                                 |
| 9       | save\_asset            | Saving assets of the Respondent                                        |
| 10      | living\_expenses       | Living costs of the Respondent                                         |
| 11      | other\_expenses        | Other expenditure of the Respondent                                    |
| 12      | incoming\_salary       | Incoming income of the Participant's                                   |
| 13      | incoming\_own\_farm    | The Participant's own incoming farm                                    |
| 14      | incoming\_business     | The Participant's incoming company                                     |
| 15      | incoming\_no\_business | Incoming intensive agricultural fees of the Respondent                 |
| 16\.    | incoming\_agricultural | Income comes from agricultural sector                                  |
| 17      | farm\_expenses         | Total expenditure in perspective of farm                               |
| 18      | labor\_primary         | Status of the labor needed in primary stages                           |
| 19      | lasting\_investment    | Total amount of lasting investment                                     |
| 20      | no\_lasting\_investmen | Total amount of no lasting investment                                  |
| 21      | target                 | \[Zero: No depressed\] or \[One: depressed\] (Binary for target class) |

**Preprocessing (Data normalization)**

For minimizing execution time and improve results, the data need to filter in the first step. We normalize the data for this reason so that the characteristics continue to follow: We used min-max function scaling (normalization) for all the aspects in this study. It is a method of re-scaling and moving aptitudes so that they end up right in the middle.

***x***<sub>normalized</sub> = (***x*** – ***x***<sub>minimum</sub>) / (***x***<sub>maximum</sub> – ***x*** <sub>minimum</sub>)

**Feature Importance Plot**

The feature value specifies that the features in the data set are more practical or significant than others. Employing feature extraction might allow you to understand better the problem you have solved and improve your model in some situations. The process of assigning a value to input data based on its effectiveness in anticipating a target variable is known as feature value.

**Machine learning Technique**

**Random Forest (RF)**

A RF is a classification and prediction technique based on data mining and machine learning. Ensemble learning is a method for solving complicated problems that incorporates several classifiers. Many decision trees are used in a random forest method. The random forest algorithm used bagging or bootstrap aggregation to generate the 'tree.' Bagging is a term that refers to the grouping of machine learning approaches to improve their accuracy. The (random forest) algorithm calculates the outcome based on tree-based predictions. It forecasts by averaging or combining the output of different trees. As the number of trees increases, the accuracy of the output improves. Using a random forest technique, you can avoid the drawbacks of a decision tree algorithm. \[22\].

**Logistic Regression (LR)**

When the response variable is categorical, LR is the best regression strategy to use (binary). The approach for determining the risk factor is usually based on the assumption that the independent variables are normally distributed with equal variances. However, in most real-world circumstances, some of the variables are qualitative or measured on nominal or ordinal scales, which violates the normalcy assumption. The logistic regression model is then used, which does not require any distributional assumptions.

Assume that there are survivors of the survival-related incident. Some are referred to as successes, while others are referred to as failures. Let , if the *i<sup>th</sup>* individual is a success and , if the *i<sup>th</sup>* individual is a failure. Suppose that for each of the individuals, independent variables are measured. These characteristics could be qualitative or quantitative in nature. The issue is determining how to link the independent variables to the dichotomous dependent variables .n the traditional regression model method assuming ’s are normally distributed with mean and variance and is the probability of success and is the probability of failure i.e.,

is linearly dependent on ’s. The model may be written as

Cox develop a model which is called linear logistic regression model and is given by

where are unknown constants.

where is called the logistic transform of and is called a linear logistic model. Another name of is log odds \[23\].

![A picture containing text, electronics Description automatically generated](625a7bb4a8830_media/media/image24.png)

**Figure-1:** Graphical view of Logistic function \[28\].

**Naive Bayes (NB)**

For very large volumes of data, the Naive Bayes (NB) classifier framework is simple to build. It's a mathematical model based on the Bayes' rule and premised on separate determinants. Of basic terms, an NB learning approach based on a certain characteristic in a class has no bearing on any other functionality. The calculation of class conditional density \[25\] is a major flaw in the naive Bayes technique. Depending on the data points, the conditional class density is usually determined. As a result, for unknown classification issues, we may be able to determine the conditional class density from unknown data objects designated by probability distributions. The equation made it possible to determine the likelihood function for each given situation \(P(c)\), \(P\left( x \middle| c \right)\) and \(P(x)\). Look beneath the equation:

where \(P\left( . \right)\ and\ P(⃓\ )\)denotes the probability and the conditional probability, respectively, \(P\left( c \middle| x \right)\ \)seems to be the posterior probability of group (target) includes integrated (attribute), \(P(c)\) is the reflection coefficient of class, \(P\left( x \middle| c \right)\) is the probability of class received indicator, and \(P(x)\) is the likelihood of determinant.

\[P\left( c \middle| x \right) = \frac{P\left( x \middle| c \right)P(c)}{P(x)}\ldots\ldots\ldots\ldots(2)\]

**Evaluation Criteria**

**Confusion matrix:**

The efficiency of a classification model is evaluated using a \(n \times n\) matrix, where n is the number of core points. The matrices produce the overall output scores based on the classifier's predictions. It provides a clear picture of where our classification method is working and what inaccuracies it generates. We would use 2 × 2 matrices with four values for a binary classifier, as seen below:

<table>
<thead>
<tr class="header">
<th></th>
<th>Actual Value</th>
<th></th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td><blockquote>
<p>Predicted Value</p>
</blockquote></td>
<td>TP</td>
<td>FP</td>
</tr>
<tr class="even">
<td></td>
<td>FN</td>
<td>TN</td>
</tr>
</tbody>
</table>

**Figure-2:** pattern of confusion matrix.

The true positive value, true negative value, false positive value, and false negative value are used to calculate the accuracy, precision, recall, specificity, and F1-score \[24\].

**True positive (TP):** the real observation suggests that depression exists, and the machine learning algorithm diagnoses depression from the given data (i.e., the detection result is true positive).

**TN (true negative):** the real observation suggests that depression exists, however the ML system is unable to detect depression based on the data provided (i.e., the detection result is a true negative).

**FN (false positive):** The real observation suggests that no depression exists, and the ML algorithm indicates that no depression is recognized from the given data .

**FN (false negative):** the real observation reveals that there is no depression, however the ML system detects depression from the available data (i.e., the detection result is a false negative).

**Accuracy**

Accuracy, precision, recall, and f1-measure are four critical variables for evaluating categorization outcomes. Accuracy is one of the most important categorization grading criteria. \[ 2\]. which is expressed as the following:

\[Accuracy = \frac{\text{TP} + \text{TN}}{TN + FN + TP + FP}\]

**Precision**

\[Precision = \frac{\text{TP}}{TP + FP}\]

What percentage of all positive instances is meant to be positive? The denominator is the model prediction that renders as positive from the entire dataset. Consider it a test to see how accurate the model is when it claims to be correct \[25\].

**Recall**

\[Recall = \frac{\text{TP}}{TP + FN}\]

Good instances are a percentage of overall positive instances; throughout this case, the number of positive occurrences in the data is the denominator. Assume we are attempting to determine then how many accurate ones the system skipped while they were available. As a result, the formula for measuring recall/sensitivity is as follows \[25\].

**F1-Score**

\[F1 - Score = 2*\left( \frac{Precision*Recall}{Precision + Recall} \right)\]

It is the amount of recall and precision in a harmonic form. This is where both factors come into play; the higher the F1 score, the better. The technique achieves well in the F1 ranking if the expected positive is actual positive (precision) and the model does not lose out on positives when predicting negatives (recall). The F1 score, which is a number between 0 and 1, represents the mean value. The F1 ranking, which is as follows, can be used to assess accuracy \[26\].

**Support:**

Support is the number of correct representations of the entity in the provided dataset. Uneven training data support could suggest serious problems with the classifier's claimed scores, necessitating stratified sampling or reorganization\[27\].

**True-positive and True-negative:**

The program accurately classified a data point as genuine or false \[ 25\].

**False-positive or false-negative:**

A false positive or false negative data item is one that the algorithm incorrectly identifies. If the system mistakenly selected an incorrect data point as accurate, it would be a false positive. \[25\].

**K-fold – cross validation technique:**

We randomly split the dataset into training and test sets while creating machine learning models, with the training set including the majority of the data. Despite the fact that the test dataset is tiny, there is still a risk that we overlooked some crucial information that could have improved the model. There's also the issue of the training set's large variation. The concept of K-fold cross-validation comes in help here. The training set is randomly divided into K (typically between 5 and 10) subsets known as folds in K-fold Cross-Validation. The model is trained using K-1 folds, then the model is tested using the other fold Because the training and test folds are chosen at random, this strategy improves the high variance problem in a dataset. \[29\].

**Proposed architecture:**

The proposed scheme depicts in Figure- 3 as a pictorial display. Data collecting, data preprocessing, data splitting, oversampling technique implementation, machine learning algorithm for model training, and model output evaluation were all part of the design.

![Diagram Description automatically generated](625a7bb4a8830_media/media/image25.jpeg)

**Figure-3**: Proposed Model Architecture

**Results and Analysis**

**Feature importance plot**

The feature importance score indicates how useful or significant each feature was in the model's growth. The function importance calculates by using the classifier model. It illustrates in Figure-4. In Figure-4, the feature "Age" , "lasting\_investment”, and "no\_lasting\_investment" were the dataset's top three essential features. Where "incoming\_salary," "sex," and "incoming\_business" are the less critical features in our dataset.

![Chart Description automatically generated](625a7bb4a8830_media/media/image26.png)

**Figure 4:** Important Feature Score

Now we will examine four assessment methods for our dataset. Table 2 and figure 5 show that with an imbalanced dataset, the random forest provided the highest accuracy 83%, where naïve Bayes provides the lowest accuracy 81%. It means that for the imbalance dataset, our three-machine learning algorithm provides the most insufficient accuracy. If we also observe that other criterion, in all of the cases highest for random forest algorithm.

**Table 2: Percentage of classification results with imbalance data**

| **Methods** | **Accuracy** | **Precision** | **Recall** | **F1** |
| ----------- | ------------ | ------------- | ---------- | ------ |
| Rf          | 0.83         | 0.9           | 0.3        | 0.45   |
| Lr          | 0.82         | 0.8           | 0.26       | 0.39   |
| Nb          | 0.81         | 0.57          | 0.19       | 0.28   |

<span class="chart">\[CHART\]</span>

**Figure-5: Percentage of classifications containing data on imbalance**

![Graphical user interface, application Description automatically generated](625a7bb4a8830_media/media/image27.png)

**Figure-6**: Confusion matrix for NB, LR, and RF algorithms before data balancing and standardization

**Figure 6** corresponds to the confusion matrix for NB, LR, and RF algorithms' unbalanced data**.** These uncertainty matrices' false-negative and true negative values are higher than the true positive and false-positive values, as seen in above figure.

There are various strategies for dealing with imbalanced datasets in machine learning algorithms, such as under-sampling, over-sampling, smote, and so on. To deal with the unbalanced data in this dataset, we used oversampling considering K-fold(10-fold) validation techniques. Until then, after enforcing oversampling, there were substantial differences in machine learning model efficiency. **Table-3** and **figure-7** provide a detailed breakdown of the overview. We connected our three models to existing methods and using oversampling for imbalanced data resulted in significantly better results—our proposed method basis on a random forest model. Our designed methodology with a random forest model surpassed most current models with the maximum accuracy of 89 percent and precision of 90 percent. Here also naïve Bayes provides the lowest accuracy and precision. If we look at other measurement criteria, random forests meet all measurement criteria and outperform.

**Table-3: Percentage of classification results with balance data**

| **Methods** | **Accuracy** | **Precision** | **Recall** | **F1** |
| ----------- | ------------ | ------------- | ---------- | ------ |
| Rf          | 0.89         | 0.9           | 0.48       | 0.63   |
| Lr          | 0.87         | 0.86          | 0.42       | 0.56   |
| Nb          | 0.86         | 0.8           | 0.36       | 0.5    |

<span class="chart">\[CHART\]</span>

**Figure-7: The proportion of categorization outcomes with balanced data (Using oversampling)**

![Graphical user interface Description automatically generated](625a7bb4a8830_media/media/image28.png)

**Figure 8:** Confusion matrix for the NB, LR, and RF algorithms after data balancing with normalization

**Figure 8** displays the confusion matrix for the NB, LR, and RF algorithms after data balancing and normalization, correspondingly. These graphs illustrate that these confusion metric’s true positive and true negative value increased compared to confusion matrix **Figure 4** of before normalization. It indicates that the overall performance increased after balancing with normalization. In addition, when compared to the unbalanced data set, the false positive and false negative values in most algorithms dropped. As a result, the value of accuracy, precision, recall, and the f-1-score is frequently increased once we reconcile our data**.**

**Conclusion**

The major goal of this study was to use the K-fold validation (10-fold) technique to evaluate the performance of three distinct machine learning categorization models. We used the Over-sampling technique to improve model efficiency within that study, we also noticed a considerable improvement in a variety of structural performance measures. Rather than machine learning classifiers, random forest classifiers can usually do well. Random forest classifier achieved maximum accuracy of 89 percent and precision of 90 percent after using oversampling, while nave Bayes achieved the lowest accuracy and precision for balanced dataset. To achieve greater efficiency of possible study outcomes oversampling strategies can be combined with machine learning technologies. According to the experimental results for imbalanced data in random forests, the highest accuracy is 83 percent, and the highest precision is 90 percent. The lowest accuracy for Naive Bayes is 81 percent for the imbalance data, whereas 86 percent for balance data. In most cases, we can see that accuracy and precision have improved after using oversampling to balance data. In a nutshell, the proposed oversampling tools with balanced data set enhance the overall performance of our model.

**References**

> \[1\] Sau, A., Bhakta, I. (2017)"Predicting anxiety and depression in elderly patients using machine learning technology. "*Healthcare Technology Letters* 4 (6)**:** 238-43

\[2\] J. Choi, J. Choi, and H. T. Jung, “Applying Machine-Learning Techniques to Build Self-reported Depression Prediction Models,” *CIN - Comput. Informatics Nurs.*, vol. 36, no. 7, pp. 317–321, 2018, doi: 10.1097/CIN.0000000000000463.\[3\] K. Kipli, A. Z. Kouzani, and I. R. A. Hamid, “Investigating Machine Learning Techniques for Detection of Depression Using Structural MRI Volumetric Features,” *Int. J. Biosci. Biochem. Bioinforma.*, vol. 3, no. 5, pp. 444–448, 2013, doi: 10.7763/ijbbb.2013.v3.252.\[4\] J. F. Dipnall *et al.*, “Fusing data mining, machine learning and traditional statistics to detect biomarkers associated with depression,” *PLoS One*, vol. 11, no. 2, pp. 1–23, 2016, doi: 10.1371/journal.pone.0148195.\[5\] F. Hasanzadeh, M. Mohebbi, and R. Rostami, “Prediction of rTMS treatment response in major depressive disorder using machine learning techniques and nonlinear features of EEG signal,” *J. Affect. Disord.*, vol. 256, no. May, pp. 132–142, 2019, doi: 10.1016/j.jad.2019.05.070.\[6\] S. Jiménez-Serrano, S. Tortajada, and J. M. García-Gómez, “A mobile health application to predict postpartum depression based on machine learning,” *Telemed. e-Health*, vol. 21, no. 7, pp. 567–574, 2015, doi: 10.1089/tmj.2014.0113.\[7\] A. Priya, S. Garg, and N. P. Tigga, “Predicting Anxiety, Depression and Stress in Modern Life using Machine Learning Algorithms,” *Procedia Comput. Sci.*, vol. 167, no. 2019, pp. 1258–1267, 2020, doi: 10.1016/j.procs.2020.03.442.\[8\] Y. Zhang, S. Wang, A. Hermann, R. Joly, and J. Pathak, “Development and validation of a machine learning algorithm for predicting the risk of postpartum depression among pregnant women,” *J. Affect. Disord.*, vol. 279, no. September 2020, pp. 1–8, 2021, doi: 10.1016/j.jad.2020.09.113.\[9\] Z. Xie, O. Nikolayeva, J. Luo, and D. Li, “Building risk prediction models for type 2 diabetes using machine learning techniques,” *Prev. Chronic Dis.*, vol. 16, no. 9, pp. 1–9, 2019, doi: 10.5888/pcd16.190109.\[10\] M. Srividya, S. Mohanavalli, and N. Bhalaji, “Behavioral Modeling for Mental Health using Machine Learning Algorithms,” *J. Med. Syst.*, vol. 42, no. 5, 2018, doi: 10.1007/s10916-018-0934-5.\[11\] T. Iliou *et al.*, “ILIOU machine learning preprocessing method for depression type prediction,” *Evol. Syst.*, vol. 10, no. 1, pp. 29–39, 2019, doi: 10.1007/s12530-017-9205-9.\[12\] S. Kumar and I. Chong, “Correlation analysis to identify the effective data in machine learning: Prediction of depressive disorder and emotion states,” *Int. J. Environ. Res. Public Health*, vol. 15, no. 12, 2018, doi: 10.3390/ijerph15122907.\[13\] B. Hosseinifard, M. H. Moradi, and R. Rostami, “Classifying depression patients and normal subjects using machine learning techniques and nonlinear features from EEG signal,” *Comput. Methods Programs Biomed.*, vol. 109, no. 3, pp. 339–345, 2013, doi: 10.1016/j.cmpb.2012.10.008.\[14\] M. R. Islam, M. A. Kabir, A. Ahmed, A. R. M. Kamal, H. Wang, and A. Ulhaq, “Depression detection from social network data using machine learning techniques,” *Heal. Inf. Sci. Syst.*, vol. 6, no. 1, pp. 1–12, 2018, doi: 10.1007/s13755-018-0046-0.\[15\] R. M. Khalil and A. Al-Jumaily, “Machine learning based prediction of depression among type 2 diabetic patients,” *Proc. 2017 12th Int. Conf. Intell. Syst. Knowl. Eng. ISKE 2017*, vol. 2018-Janua, pp. 1–5, 2017, doi: 10.1109/ISKE.2017.8258766.\[16\] I. Fatima, B. U. D. Abbasi, S. Khan, M. Al-Saeed, H. F. Ahmad, and R. Mumtaz, “Prediction of postpartum depression using machine learning techniques from social media text,” *Expert Syst.*, vol. 36, no. 4, pp. 1–13, 2019, doi: 10.1111/exsy.12409.\[17\] A. Ghandeharioun *et al.*, “Objective assessment of depressive symptoms with machine learning and wearable sensors data,” *2017 7th Int. Conf. Affect. Comput. Intell. Interact. ACII 2017*, vol. 2018-Janua, pp. 325–332, 2017, doi: 10.1109/ACII.2017.8273620.\[18\] I. Nouretdinov *et al.*, “Machine learning classification with confidence: Application of transductive conformal predictors to MRI-based diagnostic and prognostic markers in depression,” *Neuroimage*, vol. 56, no. 2, pp. 809–813, 2011, doi: 10.1016/j.neuroimage.2010.05.023.\[19\] I. Bhakta and A. Sau, “Prediction of Depression among Senior Citizens using Machine Learning Classifiers,” *Int. J. Comput. Appl.*, vol. 144, no. 7, pp. 11–16, 2016, doi: 10.5120/ijca2016910429.\[20\] A. Khodayari-Rostamabad, J. P. Reilly, G. M. Hasey, H. de Bruin, and D. J. MacCrimmon, “A machine learning approach using EEG data to predict response to SSRI treatment for major depressive disorder,” *Clin. Neurophysiol.*, vol. 124, no. 10, pp. 1975–1985, 2013, doi: 10.1016/j.clinph.2013.04.010.\[21\] N. Koutsouleris *et al.*, “Prediction Models of Functional Outcomes for Individuals in the Clinical High-Risk State for Psychosis or with Recent-Onset Depression: A Multimodal, Multisite Machine Learning Analysis,” *JAMA Psychiatry*, vol. 75, no. 11, pp. 1156–1172, 2018, doi: 10.1001/jamapsychiatry.2018.2165.\[22\] Khosrowabadi, Reza, et al. "A Brain-Computer Interface for classifying EEG correlates of chronic mental stress." *The 2011 international joint conference on neural networks*. IEEE, 2011.\[23\] Kleinbaum, D. G., Dietz, K., Gail, M., Klein, M., & Klein, M. (2002). *Logistic regression* (p. 536). New York: Springer-Verlag.\[24\] Beauxis-Aussalet, Emma, and Lynda Hardman. "Simplifying the visualization of confusion matrix." *26th Benelux Conference on Artificial Intelligence (BNAIC)*. 2014 \[25\] Davis, Jesse, and Mark Goadrich. "The relationship between Precision-Recall and ROC curves." *Proceedings of the 23rd international conference on Machine learning*. 2006.\[26\] Chicco, Davide, and Giuseppe Jurman. "The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation." *BMC genomics* 21.1 (2020): 1-1\[27\] Gunn, Steve R. "Support vector machines for classification and regression." *ISIS technical report* 14.1 (1998): 5-16.\[28\] https://towardsai.net/p/machine-learning/logistic-regression-with-mathematics\[29\] Rodriguez, Juan D., Aritz Perez, and Jose A. Lozano. "Sensitivity analysis of k-fold cross validation in prediction error estimation." *IEEE transactions on pattern analysis and machine intelligence* 32.3 (2009): 569-575.
