EXPLAINABLE MACHINE LEARNING FOR INSURANCE CLAIM FRAUD DETECTION: A COMPARATIVE PYTHON-BASED STUDY

MEMETI, Ermira and ZDRAVEVSKI, Eftim and Luma-Osmani, Shkurte and Imeri, Florinda (2026) EXPLAINABLE MACHINE LEARNING FOR INSURANCE CLAIM FRAUD DETECTION: A COMPARATIVE PYTHON-BASED STUDY. In: INTERNATIONAL CONFERENCE “FROM RESEARCH TO APPLICATION”, 20 May, 2026.

[img] Text
Proceedings Book 2026-2 (1)-508-516.pdf

Download (588kB)

Abstract

Today insurance fraud is a big problem for insurance companies since it is challenging to detect potential fraud cases within a large amount of claims data. Traditional rule-based models can detect already known fraud cases, but they often fail to detect new ones and can flag too many false positives. In this paper, several machine learning algorithms will be compared to detect insurance claim fraud. The analysis will focus on comparing models’ performance, handling class imbalances, interpretability and feature engineering. A publicly available dataset with 1,000 cases of insurance claims will be used in the analysis and at the same time, eight classification algorithms will be considered: majority-class baseline, Gaussian Naive Bayes, k-NN, weighted logistic regression, decision tree, random forest, gradient boosting, and XGBoost. To ensure responsible modeling, the study leaves out direct identifiers and sensitive variables like sex, ZIP code, education, occupation, hobbies, and relationship status. The analysis includes missing-value treatment, date-based feature engineering, one-hot encoding, stratified train-test splitting, 5-fold stratified cross-validation, and model evaluation using accuracy, precision, recall, F1-score, ROC- AUC, and confusion matrices. The results indicate that XGBoost achieved the highest cross-validated F1-score, while the decision tree achieved the strongest F1-score on the holdout test set. Random forest had the highest holdout ROC-AUC, and weighted logistic regression was a useful and interpretable benchmark. SHAP analysis found that factors like incident severity, vehicle claim amount, policy age, incident state, vehicle year, annual premium, and other claim-related variables affected the model’s predictions.

Item Type: Conference or Workshop Item (Paper)
Subjects: Q Science > QA Mathematics
Q Science > QA Mathematics > QA76 Computer software
Divisions: Faculty of Engineering, Science and Mathematics > School of Electronics and Computer Science
Depositing User: Unnamed user with email zshi@unite.edu.mk
Date Deposited: 21 Sep 2026 09:52
Last Modified: 21 Sep 2026 09:52
URI: http://eprints.unite.edu.mk/id/eprint/2419

Actions (login required)

View Item View Item