Why thin-file approvals look fine in backtest and fail in production

One narrow band of the score distribution, circled. The dashboards behind it are green, and they will stay green.
The model passed the lab. The cutoff did not survive drift, mix, or the board.
A thin-file scorecard clears validation with a Gini nobody argues with. It ships in March. A year later the approval rate is up, the loss rate looks flat, and somebody puts the chart in a board deck. Six months after that the vintages mature and the flat line bends.
Nobody shipped a broken model. It is doing in production exactly what it did in validation. The problem is that a scorecard and a cutoff are two different objects, and only one of them was tested.
The scorecard ranks. The cutoff decides. Validation measures ranking, on a population assembled by the previous cutoff, over a window long enough for the outcomes to have arrived. Production runs the cutoff, on a population the new policy is actively reshaping, and reports on outcomes that have not arrived yet. Very little of what validation proved transfers to what production is doing.
Your backtest has never met the people you are about to approve
Ask where the training labels came from. Every applicant with a repayment outcome is an applicant somebody approved. Declines have no outcome, because nothing was lent, so they are not in the data.
Your model learned one thing precisely: among people the old cutoff already liked, who repays. That is a real question and it is not the question you are asking. You are about to approve people the old cutoff declined. They sit outside the region your data covers. The model will still produce a score for them, confidently, because a model always produces a score.
The industry has a name for this, reject inference, and a shelf of techniques for patching it. The patches are assumptions wearing statistics. Every one of them has to guess what would have happened to people nothing happened to.

Labels exist only above the old cutoff. The band between the two cutoffs is the entire population the new policy adds, and nothing was ever lent there. Schematic, not measured data.
Thin-file is where the gap runs widest. A thin file is close to a definition of a file the old policy could not price, so it declined. The population you most want to expand into is the population your evidence covers worst.
Bad loans need time to go bad
Credit losses arrive late. A book written this quarter does not show its true loss rate for four to six quarters, longer on longer terms.
Run the consequence. In month three the new policy has more approvals and almost no defaults, because almost nothing has had time to default. In month nine it still looks good. The early read on a loosened cutoff comes back positive by construction, and it comes back positive whether the policy was sound or reckless. The two are indistinguishable for about a year.

The expansion decision lands inside the window where almost nothing has had time to default. A sound policy and a reckless one look identical here. Schematic, not measured data.
The approval rate, meanwhile, is visible daily. The signal that arrives fast is the one the policy was meant to move. The signal that arrives slow is the one that says whether it worked. An organization will act on the fast one, expand the policy, and price the expansion before the first honest number lands.
Vintage curves fix this and most lenders have them. The failure is not an absent tool. It is that the meeting where a policy gets expanded runs on the quarterly cycle, and the vintage does not.
You watch rank order. The cutoff spends calibration.
A scorecard has two properties that drift separately.
Rank order is whether the model sorts riskier applicants below safer ones. Gini and AUC measure it, the monitoring pack reports it, and it is remarkably durable. A scorecard can hold its rank order for years.
Calibration is whether a given score still carries the default probability it carried when the cutoff was set. It drifts with the economy, with channel mix, with anything that moves the base rate.

Rank order is durable and it is what gets reported. Calibration is fragile and it is what the threshold actually spends. Both series sit on an arbitrary shared scale; only the shapes carry meaning.
The cutoff consumes calibration. Nobody approves a percentile. They approve everyone above a number, and that number was chosen because it corresponded to a loss rate the business would accept. When calibration slips while rank order holds, every dashboard stays green and the cutoff quietly starts buying a different risk than the one it was priced for.
The policy that ships is the cutoff plus an override layer
The modeled policy is a threshold. The deployed policy is a threshold plus whatever the humans do around it.
Overrides, exception queues, relationship approvals, the branch that has a reason. These land almost entirely on the margin, because nobody overrides an obvious decline or an obvious approve. They cluster exactly where the cutoff lives, which is exactly where the model is least certain and the backtest is thinnest.
That layer is rarely modeled and almost never remeasured. It is a second policy running in production, with no owner and no validation.
What remeasuring thin-file approvals actually means
A better model will not fix this. A better model sharpens rank order, which is already the healthy part. The cutoff is the object under test, and it can be measured as one.
Four questions, asked per segment rather than per portfolio:
What share of the population the new cutoff approves has any analogue in the training data? Not the average applicant. The marginal one.
What does the vintage curve say at matched months on book, against the policy it replaced?
When was the score-to-probability mapping last refit, as opposed to the model last rebuilt?
What share of decisions near the cutoff is set by override, and who reviews that layer?
The last one is usually the fastest to answer and the most uncomfortable.
A model that passes validation has proved it can rank the past. A cutoff has to survive a population it has not met, on evidence that has not arrived. Those are different tests, and only one of them was run.
If your approval rate moved this quarter and your oldest affected vintage is younger than a year, you do not yet know whether the policy worked. You know the approval rate moved.




Comments