Machine Learning / Research Project / May 2025 - Aug 2025

Credit Risk Model Replication

Replication study comparing Logistic Regression, Random Forest, and PLTR on 30,000 credit-card clients to test the trade-off between predictive performance and explainability.

Role
Co-author / Modeling and research lead
Scope
Random Forest, PLTR, experiment design, report and figures
Dataset
30,000 clients / six-month payment histories
30,000
borrower records
79%
highest accuracy
0.771
best replication AUC
6 months
behavioral history

Credit model decision

Performance matters. So does explainability.

6,000 held-out observations
Model benchmarkHeld-out results
Logistic Regression77% accuracy
AUC 0.771Best AUC, globally interpretable
Random Forest79% accuracy
AUC 0.765Highest accuracy, lower transparency
PLTR75% accuracy
AUC 0.712Local explanations, replication-sensitive
Feature progressionLogit accuracy
Demographics49%
+ Financial behavior64%
+ Payment delay70%
+ Credit usage77%

Behavioral payment signals added more predictive value than static demographics.

Decision: Random Forest led on accuracy; Logistic Regression offered the strongest balance of discrimination and transparency.

The question

Credit risk models need to distinguish likely defaults while remaining understandable enough for validation and governance. This study recreates a published credit-scoring comparison under a controlled feature set and asks whether a hybrid model can improve non-linear prediction without giving up interpretability.

Pythonpandasstatsmodelsscikit-learnLogistic RegressionRandom ForestPLTR

Research report

Credit Modeling Research Project