Background¶
With new credit score models on the rise, and with the GSE's release of these scores for mortgage originations going back to 2013, there has been ongoing analysis of the characteristics and performance of new credit score models versus classic scores, as well as the ramifications of potential score inflation under lender choice. However, little attention has been paid to the possibility that some loans may have multiple scores reported; in this case, theoretically, the use of multiple scores could provide greater accuracy in predicting mortgage delinquency. This analysis takes a closer look at the question of how to handle loans that have reported both a Classic FICO (CF) and VantageScore 4.0 (VS4) score. We use historical data to find the score-based model that most accurately predicts transitions from current to delinquent (CtoD) by using both scores.
In order to predict whether loans will go delinquent or not, we develop models with varying methods of combining the reported scores. The metrics used to compare these models include bias, ROC/AUC, Kolmogorov–Smirnov (K-S) statistic, Brier score, Maximum Vertical Distance (MVD), and log loss. Visualizations are also used to better understand how these models align with the actual data.
Data¶
The models are based on loan-level data from the Fannie Mae database. The purpose of this analysis is to evaluate loans containing both a CF score and a VS4 score; therefore, the data only contains such loans. The charts below show the data availability of CF and VS4 scores for mortgages originated between 2012 and 2023.
In addition, since our models predict the transition rate from Current to Delinquent, the model dataset consists only of observations where the loan was current at the given month, along with a field that indicates whether the loan transitioned to delinquent the subsequent month. The final dataset consists of approximately 11.9 million loan observations.
Note: Some models and tests require additional processing of the data. These specific instances will be briefly mentioned in the upcoming sections of this notebook.
Regressions¶
Our models are produced using logistic regressions. These regressions produce intercepts and slopes based on each model’s respective credit scores. These intercepts and slopes then create predicted delinquency rates that align with the data.
The logistic regression is done using the generalized linear model function. The logistic regression uses binary data; in this case, that would mean whether a loan went delinquent (1) or stayed current (0). Therefore, the regression uses the loan-level data rather than the grouped data to take in the recorded credit score(s) as its input. For this type of regression, a log odds total is produced using the intercepts and slopes, which is then algebraically manipulated to produce a predicted CtoD rate for each loan.
This type of regression makes sense for the type of data we are working with and allows us to obtain accurate results compared to a simpler model such as a linear regression.
Models¶
In this section, you will find the various models created for this study. There are five models that predict CtoD rates using different variables. All these models are tested using various metrics to determine which one provides the best interpretation of the given CF and VS4 scores on each loan.
Model 1: Only Classic FICO¶
Model 1 takes in a CF score as its only input. It enables cross-comparison between these models on a credit score basis. The logistic regression model is defined as:
$$ P(\mathrm{CtoD})= \frac{1} {1+e^{-(\alpha^{M1}+\beta^{M1}\cdot \mathrm{CS})}} $$
where:
$P(\mathrm{CtoD})$ is the predicted probability of delinquency.
$\alpha^{M1}$ is the Model 1 intercept.
$\beta^{M1}$ is the Model 1 slope.
$\mathrm{CS}$ denotes the CF credit score.
Below are the coefficients:
Code
slope_m1 = -0.017684171
intercept_m1 = 6.458873958
Model 2: Only VantageScore 4.0¶
In this model, the input for the logistic regression is simply the VS4 score attached to each loan. This model determines how effective VS4 is at predicting whether a loan will go delinquent or stay current. In this case, the logistic regression model is defined as:
$$ P(\mathrm{CtoD})= \frac{1} {1+e^{-(\alpha^{M2}+\beta^{M2}\cdot\mathrm{CS})}} $$
where:
$P(\mathrm{CtoD})$ is the predicted probability of delinquency.
$\alpha^{M2}$ is the Model 2 intercept.
$\beta^{M2}$ is the Model 2 slope.
$\mathrm{CS}$ denotes the VS4 credit score.
Below are the coefficients:
Code
slope_m2 = -0.014448933
intercept_m2 = 4.146557719
Model 3: Average of Classic FICO and VantageScore 4.0¶
Model 3 takes in the direct average of the CF and VS4 scores as its input. This model considers whether averaging the CF and VS4 scores associated with a loan, instead of using just one or the other, is helpful in predicting that loan’s delinquency rate. The logistic regression model is defined as:
$$ P(\mathrm{CtoD})= \frac{1} {1+e^{-(\alpha^{M3}+\beta^{M3}\cdot\mathrm{CS})}} $$
where:
$P(\mathrm{CtoD})$ is the predicted probability of delinquency.
$\alpha^{M3}$ is the Model 3 intercept.
$\beta^{M3}$ is the Model 3 slope.
$\mathrm{CS}$ denotes the rounded simple average of the CF and VS4 credit scores, as calculated below:
Code
df['m3_cs'] = round((df['fico_cs'] + df['vs4_cs'])/2)
Below are the coefficients:
Code
slope_m3 = -0.018575326
intercept_m3 = 7.15627315
Model 4: Average of Implied CtoD Rates¶
Model 4 is simply the average implied (i.e., modeled) CtoD rate of Models 1 and 2; it does not involve a separate regression. The intercepts and slopes from Models 1 and 2 are used to produce separate log odds, as shown below:
Code
m1_log_odds = intercept_m1 + slope_m1 * df['fico_cs']
Code
m2_log_odds = intercept_m2 + slope_m2 * df['vs4_cs']
These log odds are then used as inputs in a formula to find the average implied CtoD rates:
$$ P(\mathrm{CtoD}) = \frac{1}{2} \left( \frac{1}{1+e^{-m1_{\mathrm{log\ odds}}}} + \frac{1}{1+e^{-m2_{\mathrm{log\ odds}}}} \right) $$
where:
- $P(\mathrm{CtoD})$ is the predicted probability of delinquency, denoted by $\mathrm{m4\_icd}$ in the cell below.
These CtoD rates are then used to find the log odds for Model 4, which are converted into credit scores to be used in various test metrics and visualizations (note that this conversion is order preserving).
Code
m4_log_odds = np.log(df['m4_icd'] / (1 - df['m4_icd']))
Code
df['m4_cs'] = round(1 / slope_m2 * (m4_log_odds - intercept_m2))
Model 5: Two-Variable Model¶
Model 5 is a two-variable logistic model that considers both CF and VS4 scores together to predict CtoD rates. This model is characterized by one intercept, one slope for the CF variable, and one slope for the VS4 variable. This model is defined by:
$$ P(\mathrm{CtoD}) = \frac{1} {1 + e^{-\left(\alpha^{M5} + \beta^{M5}_{CF}\cdot CF + \beta^{M5}_{VS4}\cdot VS4\right)}} $$
where:
$P(\mathrm{CtoD})$ is the predicted probability of delinquency.
$\alpha^{M5}$ is the Model 5 intercept.
$\beta^{M5}_{CF}$ is the Model 5 slope for CF.
$\mathrm{CF}$ is the CF score.
$\beta^{M5}_{VS4}$ is the Model 5 slope for VS4.
$\mathrm{VS4}$ is the VS4 score.
This model is the mathematical equivalent of inputting a weighted average of the CF and VS4 scores. This can be thought of as the more accurate version of Model 3, which simply takes the unweighted average of the CF and VS4 scores. Below is the formula for this weighted average credit score:
Code
df['m5_cs'] = round((fico_slope_m5*df['fico_cs'] + vs4_slope_m5*df['vs4_cs'])/(fico_slope_m5 + vs4_slope_m5))
Below are the coefficients:
Code
fico_slope_m5 = -0.010542491
vs4_slope_m5 = -0.008270896
intercept_m5 = 7.326722073
Statistical Tests¶
This section summarizes the statistical metrics used to compare the models listed above. Each subsection describes a specific test and presents a table of the corresponding metrics. The section concludes with a comprehensive summary table containing all the metrics. For all tables, bold values indicate the best-performing model(s) across all tests.
Note: We use the terms 'good' and 'bad' when referring to the transitions of a (current) loan to either current or delinquent; however, these terms do not reflect our views on borrowers and are only used as shorthand classifications for our analysis.
Bias¶
The Bias test is used to verify that the models accurately capture the overall CtoD rate within the model training set. The bias is calculated by comparing the average CtoD rates of the model and actual data. Therefore, to get a bias score, the average model CtoD rate is divided by the average actual CtoD rate. A bias of 1 represents a perfect model; if a score is close to one, the model is functioning as it should.
Code
avg_actual = df[cd].mean()
avg_model = df[m].mean()
bias = avg_model/avg_actual
Table T1. Bias Model Performance Comparison
| Model 1 | Model 2 | Model 3 | Model 4 | Model 5 |
|---|---|---|---|---|
| 1.000000 | 1.000000 | 1.000000 | 1.000000 | 1.000000 |
All the bias scores produced by each model are very close to 1, so their differences are negligible. Therefore, all the models are functioning as intended.
ROC/AUC¶
To determine how effective our models are at separating the transitions to delinquency from loans that stay current, we use the ROC and AUC tests. The ROC (Receiver Operating Characteristic) involves creating a curve that depicts the cumulative percentage of all "bad" (current-to-delinquent) transitions that occur, versus the cumulative percentage of all "good" (current-to-current) transitions, as the credit score increases. The AUC (Area Under the Curve) puts this visualization into a numerical value, that allows us to better compare the various models. The best possible AUC value a model could get is 1 (Perfect Model). Therefore, a higher AUC depicts a more effective model.
Before calculating, the set of observations is grouped by credit score (using the average and weighted average in models 3 and 5, respectively, or a translation from predicted CtoD rate to credit score for Model 4). To get the AUC values, we first create a ROC curve. This curve is created by calculating the cumulative sum of good observations and the cumulative sum of bad observations as credit score increases.
Code
observed_bad = sum(roc_df['bad_mm'])
roc_df['cum_bad'] = roc_df['bad_mm'].cumsum()/observed_bad
observed_good = sum(roc_df['good_mm'])
roc_df['cum_good'] = roc_df['good_mm'].cumsum()/observed_good
The x-axis for this curve represents the percentage of all good transitions and the y-axis represents the percentage of all bad transitions.
Then, the AUC is calculated by applying the trapezoidal rule to the ROC curve.
Code
AUC = np.trapezoid(roc_df['cum_bad'], roc_df['cum_good'])
Below are the graphs showing the ROC curves which depict the separation of good and bad observations:
Table T2. AUC Model Performance Comparison
| Model 1 | Model 2 | Model 3 | Model 4 | Model 5 |
|---|---|---|---|---|
| 0.749455 | 0.745536 | 0.762389 | 0.760899 | 0.762765 |
Based on the AUC results, it can be concluded that Model 5 is best at differentiating the current and delinquent transitions, as it gives the highest AUC of 0.762765. Note that Models 3-5 all outperform single score models.
K-S Statistic
Much like the ROC and AUC tests, the Kolmogorov-Smirnov (K-S) statistic is another method of measuring separation between transitions to delinquency vs. staying current. More specifically, it is the maximum difference between the cumulative percentage of all good transitions ($\mathrm{cum\_good}$) and the cumulative percentage of all bad transitions ($\mathrm{cum\_bad}$) as the model CtoD rate ranges from highest to lowest. Calculating this metric involves grouping the dataset by the model's predicted CtoD rates and counting the number of good observations ($\mathrm{good\_mm}$) and bad observations ($\mathrm{bad\_mm}$) across these buckets. The following code is then used to calculate the K-S statistic.
Code
ks_stat_df['good_mm'] = ks_stat_df['all_mm'] - ks_stat_df['bad_mm']
observed_bad = sum(ks_stat_df['bad_mm'])
ks_stat_df['cum_bad'] = ks_stat_df['bad_mm'].cumsum()/observed_bad
observed_good = sum(ks_stat_df['good_mm'])
ks_stat_df['cum_good'] = ks_stat_df['good_mm'].cumsum()/observed_good
ks_stat_df['diff'] = abs(ks_stat_df['cum_good'] - ks_stat_df['cum_bad'])
ks_statistic = ks_stat_df['diff'].max()
K-S statistic values range between 0 and 1, where 1 indicates that the model perfectly separates transitions by delinquency as it moves along its respective credit score scale. The higher the K-S statistic is, the better the model is at separation. Below are the K-S statistics across all the models:
Table T3. K-S Statistic Model Performance Comparison
| Model 1 | Model 2 | Model 3 | Model 4 | Model 5 |
|---|---|---|---|---|
| 0.382527 | 0.372123 | 0.398022 | 0.396384 | 0.400757 |
According to these results, Model 5 produces the highest K-S Statistic. Again, Models 3-5 all outperform Models 1-2.
Brier Score¶
The Brier score considers how close the model CtoD rates are to the actual data by calculating the mean squared difference between both. Note that this comparison is done at the observation level rather than at the aggregated level. The model outcomes are therefore compared to observed outcomes of each loan, where 0 is current-to-current, and 1 is current-to-delinquent. Below is the Brier score calculation:
Code
brier_score = (((df['model_cd_rate'] - df['observed_outcome']) ** 2).mean().sum())
The Brier score ranges between 0 (total accuracy) and 1 (total inaccuracy). Below are the Brier scores across all the models.
Table T5. Brier Score Model Performance Comparison
| Model 1 | Model 2 | Model 3 | Model 4 | Model 5 |
|---|---|---|---|---|
| 0.00141116 | 0.00141130 | 0.00141074 | 0.00141076 | 0.00141073 |
As observed with the previous metrics, Models 3-5 outperform Models 1-2, with Model 5 having a slight edge over the others.
MVD¶
The Maximum Vertical Difference (MVD) provides another way of understanding the ROC Curve, but with a focus on accuracy instead of separation. It allows us to visualize how close the model predictions are to the actual data by creating separate cumulative distribution curves for each; thus, each curve represents the cumulative (predicted or actual) percent of all "bad" transitions (bad_mm) in comparison to the cumulative percent of all observations (all_mm), as credit score increases. In the process of creating this score, the bias, which was previously calculated, is removed; however, removal of bias has a negligible effect for these 5 models. To find the MVD score, we calculate the maximum vertical difference between the model and actual curves. In general, a smaller MVD implies a closer fit between the model and the actual, and a score of 0 would indicate that the model exactly matches the results of the actual data.
For the processing of the MVD test metric, we group all the observations by their Model CtoD rates and calculate the actual delinquency rate for each bucket. The credit score groups are weighted according to the number of loans in each bin.
For our interpretation of the MVD, we use a similar process as the calculation for ROC/AUC. Therefore, we calculate the cumulative number of observations, cumulative actual transitions to delinquent, and the cumulative predicted transitions to delinquent. Each of these cumulative totals are calculated separately with the cumulative model calculation requiring some extra steps.
The cumulative percent of the total is calculated by dividing the cumulative sum of observations as the model CtoD rate decreases by the sum of all observations.
Code
mvd_df['cum_total'] = mvd_df['all_mm'].cumsum()/mvd_df['all_mm'].sum()
The cumulative percent of delinquent transitions using the actual data are calculated the same way but using only the delinquent transitions instead of all observations.
Code
mvd_df['cum_actual'] = mvd_df['bad_mm'].cumsum()/mvd_df['bad_mm'].sum()
For the model cumulative, firstly, a predicted number of CtoD observations is calculated for each bucket using the model CtoD rates and multiplying them by the number of observations in each bucket. Then, the bias calculated previously for each corresponding model is used to remove any bias from these predictions. Lastly, the cumulative sum of predicted CtoD observations for the model data is calculated using the same method as before.
Code
mvd_df['modeled_cd'] = mvd_df['all_mm'] * mvd_df['model_cd_rate']
mvd_df['modeled_cd_adj_bias'] = mvd_df['modeled_cd']/bias
mvd_df['cum_model'] = mvd_df['modeled_cd_adj_bias'].cumsum()/mvd_df['bad_mm'].sum()
Finally, the absolute value of the differences between the cumulative model data and the cumulative actual data is calculated for each credit score bin. The MVD is then determined by taking the maximum difference between the actual curve and the model curve.
Code
mvd_df['abs_val_diff'] = abs(mvd_df['cum_model'] - mvd_df['cum_actual'])
MVD = max(mvd_df['abs_val_diff'])
Below are the MVD graphs which visualize how close the model data is to the actual data:
Table T6. MVD Model Performance Comparison
| Model 1 | Model 2 | Model 3 | Model 4 | Model 6 |
|---|---|---|---|---|
| 0.031605 | 0.025492 | 0.021865 | 0.058565 | 0.024656 |
Model 3 gives the best result of 0.021865 for this test, with the lowest maximum difference between the model curve and the actual curve. Meanwhile, Model 4, despite having better separation properties than Models 1 and 2, is less accurate than either according to this metric.
Log Loss¶
The log loss is used to determine how confidently our models predict the delinquency transitions, alongside their accuracy. The log loss rewards confident and accurate predictions but punishes uncertainty and incorrect predictions. The log loss for every loan observation is summed up to create a total value. A lower total indicates a better model in this case.
To calculate the log loss values, the actual outcome of the transition (1 for CtoD and 0 for CtoC) is multiplied by the natural log of the modeled CtoD probability. This value is then added to the opposing outcome multiplied by the natural log of 1 minus the probability of CtoD. The log loss is the absolute value of the result.
Code
log_loss = abs((-df[cd] * (np.log(df[m]) - (1 - df[cd])) * np.log(1-df[m])).sum())
Table T7. Log Loss Model Performance Comparison
| Model 1 | Model 2 | Model 3 | Model 4 | Model 5 |
|---|---|---|---|---|
| 120,266 | 120,468 | 119,273 | 121,225 | 119,253 |
Model 5 gives the best result for the log loss calculation, as it has the lowest log loss total of 119,253. Therefore, this model is the most confident in its results and provides the most accurate results.
Summary Table¶
Below is a summary table containing the metrics across all models. Model 5, the two-variable model, performs best across more metrics (four out of six) compared to the other models.
Table T8. Model Performance Comparison
| Metric | Model 1 | Model 2 | Model 3 | Model 4 | Model 5 |
|---|---|---|---|---|---|
| Bias | 1.000000 | 1.000000 | 1.000000 | 1.000000 | 1.000000 |
| AUC | 0.749455 | 0.745536 | 0.762389 | 0.760899 | 0.762765 |
| KS Statistic | 0.382527 | 0.372123 | 0.398022 | 0.396384 | 0.400757 |
| Brier Score | 0.00141116 | 0.00141130 | 0.00141074 | 0.00141076 | 0.00141073 |
| MVD | 0.031605 | 0.025492 | 0.021865 | 0.058565 | 0.024656 |
| Log Loss | 120,266 | 120,468 | 119,273 | 121,225 | 119,253 |
Model vs. Actual¶
For visualization purposes, we use simple charts, which compare the model predictions to the actual data. This is done by creating line graphs showing the correlation between credit scores and delinquency rates. In this case, we organize the credit scores into 20-point buckets* to make the visualization cleaner. The number of loan observations is also depicted on the chart using a standard bar graph format. This makes the distribution of credit scores clear.
**The lower and higher credit scores are binned into larger buckets to account for low amounts of data, in terms of loans or delinquencies.*
Model 1 uses the CF credit scores for the x-axis.
This graph shows that Model 1 is both overpredicting and underpredicting the data, but overall, Model 1 produces similar results to the actual data. There is more separation between the model and actual data at the ends of the x-axis due to there being very few loans in the 550-620 credit score bin and barely any delinquencies in the 800-851 bin. Therefore, the model is overpredicting delinquency rates for lower credit scores and higher credit scores. Note that at the lower end, selection bias is playing a role since the CF score was used to qualify these loans.
Model 2 uses the VS4 credit scores for the x-axis.
This chart shows that Model 2 is mostly underpredicting delinquency rates. It also overpredicts the data at the higher and lower ends of the credit score scale; however, not as much as Model 1, due to the number of loans being more spread out among the bins.
Model 3 uses the simple average of CF and VS4 (CF+VS4/2) for the x-axis.
If you were comparing the models only using these graphs, then Model 3 would be the most successful. This graph shows the model data being almost identical to the actual data between 640 and 800. The spread of loans represents a middle ground between the Model 1 and Model 2 charts; therefore, there is still some overpredicting at the ends of the credit score scale.
Model 4 uses a converted credit score based on the average of the implied CtoD rates calculated from Models 1 and 2.
This chart illustrates how Model 4 is mostly underpredicting the data, even for the lower credit scores. Therefore, this model is able to understand the trend of delinquencies in lower score bins. However, it is still overpredicting at the higher end of the credit score scale.
Model 5 uses mapped credit scores for the x-axis. These mapped credit scores represent the weighted average of CF and VS4.
This chart shows that the Model 5 predictions are very similar to the actual data. There is slight underprediction throughout and the same overprediction at the ends as seen in the previous models.
Heatmaps¶
To further compare the models, we calculate the average model outputs of every observed 10-bin CF and VS4 combination for each model. Within each bin, the average of each granular credit score combination is weighted based on the number of observations with that combination. To take this a step further, the average model CtoD rates for each bin are then mapped to a CF score scale by inverting the Model 1 equation. Below is the formula used for this conversion:
$$\text{CF}_{\text{inverted}} = -\left\lfloor \frac{\ln\left(\frac{1}{P_M} - 1\right) + \alpha_{\text{M1}}}{\beta_{\text{M1}}} \right\rfloor$$
where:
$P_m$ is the average predicted CtoD rate for the given model.
$\alpha_{\text{M1}}$ is the fitted intercept from Model 1 (the FICO-only baseline).
$\beta_{\text{M1}}$ is the fitted slope coefficient from Model 1.
Translating these probabilities into a standardized CF score scale allows us to compare all models against a single, interpretable benchmark. Using these inverted scores, heatmaps are generated for every model to visualize the inverted CF score for each pairing of CF and VS4 scores. Since these heatmaps are based on actual combinations of both scores, only score pairings containing at least 1,000 loans are included to ensure statistically meaningful analysis.
The actual data is also visualized via a heatmap by using the same processing method described above. Like the model heatmaps, this chart is populated with the CF-equivalent of a CtoD rate associated with every observed CF and VS4 score combination. However, rather than representing the modeled prediction for each score pairing, the CF scores in this heatmap represent the actual CtoD rates.
Below is the actual heatmap (Fig H0a) as well as the model heatmaps (Figs H1- H5). Darker shades indicate higher equivalent CF scores, while lighter shades indicate lower scores.
The actual heatmap from Figure H0a can be rearranged so that the vertical axis organizes borrowers by their CF score, while the top horizontal axis represents the VS4 score band minus the CF score band. This formatting allows for easy visualization of how the actual CtoD rates change as the VS4 score deviates from the CF score, as seen in Figure H0b below:
To understand how combining both scores compares to the standalone CF model, we examine Model 5 (Fig H5), the best-performing model, against the CF-only Model 1 (Fig H1).
Consider a borrower with a low CF score (670-679) and a high VS4 score (780-789). Since Model 1 strictly evaluates this borrower based on their CF score, the model overlooks the positive credit signal from their high VS4 score. Therefore, Figure H1 shows that the borrower would be assigned an average CF score of 675. This is 37 points lower than the score of 712 assigned to this borrower in the actual heatmap (Fig H0a). The actual borrower data shows that the high VS4 score provides additional information, but the CF-only model does not capture this nuance.
Meanwhile, Model 5, the two-variable model, accounts for the signal from both scores. Using the previous example, the average predicted CtoD rate that falls within that combination of scores is equivalent to a 720 CF score on the H5 heatmap: 45 points above the CF score alone.
As demonstrated above, combining both scores can be important when modeling extreme scenarios, like when a borrower has a low CF score and a high VS4 score. However, a borrower’s CF and VS4 scores are not always extremely different. The predictions assigned by Models 1 and 5 might not be so disparate in these scenarios. In any case, it is helpful to visualize the entire spectrum of the CF scores assigned by Models 1 and 5 to see where the models differ, as shown in Figure DH1 below:
This heatmap subtracts the Model 1 heatmap scores from the Model 5 scores to quantify the score difference between the two models. The horizontal axis is the VS4 score minus the CF score, and the vertical axis is the CF score.
The map uses a diverging color gradient to illustrate both the directions and magnitude of the model score differences. White regions indicate areas of convergence where Model 1 and Model 5 yield nearly identical scores. Moving away from this area, darker shades of blue mark segments where Model 5 assigns a lower score than Model 1, whereas red zones represent cells where Model 5 assigns a higher score. The intensity of the colors reflects the degree of divergence between the two models.
Figure DH1 suggests how to handle cases where both CF and VS4 scores are present. Model 1 is a CF-only model and does not consider the presence of a VS4 score. Therefore, by subtracting Model 1 from Model 5, this heatmap indicates what changes should be made to a CF score if a VS4 score is present. For example, the cell at the low CF score (670-679) and the high VS4 score (+110) suggests that if both scores are available, the CF score should, on average, be increased by 45 points to reflect the higher VS4 score. Ultimately, this chart illustrates that when VS4 is present, the CF score should be adjusted in accordance with the values presented.
To better show the impact of Model 5, the figure below visualizes how Model 5’s weights influence how the two credit scores are interpreted. This specifically shows how Model 5 pulls the credit score away from the simple average and towards the given CF score.
This heatmap depicts the relationship between the credit scores produced by the Model 5 coefficients and the Model 3 coefficients, for every CF and VS4 credit score combination. This graph again plots the CF score on the left and the difference between the VS4 and CF scores at the top. The darker the color, the more the score leans towards CF. When zooming in on the same cell as mentioned for DH1 (670-679 CF and +110 score difference), it can be said that the lower CF score has more pull on the interpreted score, as the interpreted score is 7 less than the simple average.
Discussion¶
Based on the data compiled from these tests, Model 5 is the most efficient model. This means that producing a weighted average using the two credit scores reported on a loan will provide us with a score that more accurately represents its risk of going delinquent.
Conclusion¶
The primary purpose of this analysis was to determine the best method for predicting CtoD rates when handling borrowers with both a Classic FICO score and a VantageScore 4.0. Various approaches for combining both scores were tested, ranging from baseline score models to dual-score models. Across the test metrics evaluation, Model 5 outperforms competing models across four out of six metrics. Specifically, Model 5 achieves the highest separation capability with an AUC of 0.762765 and a KS Statistic of 0.400757, alongside the lowest log loss at 119,253. Beyond the statistical tests, Model 5 also produces one of the best Actual vs. Model graphs.
The performance of Model 5 led to further analysis of this model via the heatmaps illustrated in Figures DH1 and DH2. Figure DH1 particularly provides an actionable takeaway for handling borrowers with both scores. The values that populate this chart suggest exactly where Classic FICO alone should be adjusted to account for the VantageScore 4.0 risk signal.
Moving forward, we can use this methodology to also look at additional credit scores. This could be especially relevant as FICO 10T or other new scores are introduced. Nonetheless, this analysis provides an essential baseline for modeling the event wherein both a Classic FICO score and a VantageScore 4.0 are present.
Disclaimer: © 2026 Andrew Davidson & Co., Inc. All rights reserved. Andrew Davidson & Co, Inc. has published research papers in collaboration with other firms on topics similar to those described in this publication. You must receive permission from marketing@ad-co.com prior to copying, displaying, distributing, publishing, reproducing, or retransmitting any of the content contained in this whitepaper.
This publication is believed to be reliable, but its accuracy, completeness, timeliness, and suitability for any purpose are not guaranteed. All opinions are subject to change without notice. Nothing in this publication constitutes (1) investment, legal, accounting, tax, or other professional advice or (2) any recommendation or solicitation to purchase, hold, sell, or otherwise deal in any investment. This publication has been prepared for general informational purposes, without consideration of the circumstances or objectives of any particular investor. Any reliance on the contents of this publication is at the reader’s sole risk. All investment is subject to numerous risks, known and unknown. Past performance is no guarantee of future results. For investment advice, seek a qualified investment professional.