A comparative study of Machine Learning algorithms for early prediction of heart disease
A comparative study of Machine Learning algorithms for early prediction of heart disease
JAMIA HAMDARD
(Deemed to be University)
New Delhi 110062,
India Centre for Distance and Online Education
Centre for Distance and Online Education, Jamia Hamdard for the partial fulfillment of the degree of MCA
MISBAH KHAN
Programme: MCA
Under the guidance of:
Dr. Abdul Majid Farooqi
Assistant Professor
Centre for Distance and Online Education, Dept. of CSE Jamia Hamdard, New Delhi, India
Cardiovascular disease (CVD) remains one of the most serious health challenges worldwide and continues to be the leading cause of death. Millions of people lose their lives each year due to heart related illness, making early detection and prevention extremely important. Accurate prediction of heart disease can help healthcare professionals identify high risk patients at the early stage and provide timely treatment, ultimately improving patient outcomes. In recent years, machine learning techniques have shown significant potential in analyzing medical data and supporting clinical decision-making.
This study presents a comparative analysis of nine machine learning algorithms for the early prediction of heart disease, including Logistic Regression, Naïve Bayes, K-Nearest Neighbors, Decision Tree, Support Vector Machine, Random Forest, Gradient Boosting, XGBoost, and Artificial Neural Networks. The research utilises the UCI Cleveland Heart Disease dataset along with additional heart disease datasets to improve the reliability and robustness of the analysis.
Before model training, several preprocessing techniques were applied, including handling missing values, feature scaling, balancing class distribution using SMOTE, and selecting the most relevant features through chi-square analysis. The performance of each model was evaluated using ten-fold cross-validation and multiple evaluation metrics such as accuracy, precision, recall, F1-score, specificity, and AUC-ROC. To enhance model interpretability, SHAP analysis was employed to identify the factors that contribute most significantly to prediction outcomes.
The experimental results indicate that XGBoost achieved the best overall performance, obtaining an accuracy of 90.4% and an AUC-ROC score of 0.94. Random Forest and Gradient Boosting also demonstrated strong predictive capabilities, while traditional algorithms such as Logistic Regression and SVM delivered comparatively lower accuracy. Feature importance analysis revealed that chest pain type, maximum heart rate achieved, ST depression, number of major vessels, and age were among the most influential predictors of heart disease.
The findings of this study demonstrate the effectiveness of machine learning techniques in cardiovascular risk prediction and highlight the potential of advanced ensemble models for supporting healthcare professionals in early diagnosis and clinical decision-making.
Keywords: Heart Disease Prediction, Machine Learning, XGBoost, Random Forest, Support Vector Machine, Logistic Regression, AUC-ROC, SHAP, Explainable AI, Cleveland Dataset, Cardiovascular Disease, Clinical Decision Support.