Skip to main content
Compliance

Fair lending analysis in practice

The hard parts of this work are not statistical. You are usually not permitted to collect the variable the analysis is about, the controls that seem most natural are the ones that destroy the result, and the question a reviewer cares most about is one most lenders never document at all.

Engineering perspective, not legal advice This describes how the analysis is built and where it goes wrong technically. Fair lending is a legal field with an enforcement posture that has changed more than once in recent years through rulemaking, litigation and executive action. Nothing here is a statement of what the law currently requires of you. Read the statutes and regulations at their source, and let counsel set policy before any of this is run — including whether and how to run it at all.

Intent is not the question

The first thing to understand about fair lending analysis is that most of it has nothing to do with what anyone meant. The Equal Credit Opportunity Act and its implementing regulation prohibit discrimination in any aspect of a credit transaction on a list of bases — race, color, religion, national origin, sex, marital status, age, the fact that income comes from public assistance, and the good-faith exercise of consumer credit rights. The Fair Housing Act adds its own protected classes for housing-related credit. Neither asks whether the lender was trying.

A well-run lender can produce a disparity through a policy nobody thought about twice: a minimum loan amount, a branch footprint inherited from an acquisition, a discretionary pricing band, a data source that behaves differently across populations. Finding those before somebody else does is the entire purpose of the analysis, and the mindset that makes it work is closer to quality engineering than to legal defense.

You are probably here because

  • You are deploying a machine-learned underwriting or pricing model for the first time
  • An examination, an investor or a bank partner asked for your fair lending testing
  • Your portfolio shows a disparity and nobody can say whether it is real or noise
  • You do not collect race or ethnicity and are not sure how the analysis is even possible

The proxy section explains what is possible without the data. The control variable section is where most in-house analyses quietly go wrong. The alternative search is the part reviewers ask about and almost nobody documents.

Three theories, and only one is about a decision-maker

Practitioners work with three analytical frames, and mixing them up leads to testing for the wrong thing.

Overt discrimination is a policy or statement that treats a protected class differently on its face. It is rare, it is found by reading documents and listening to how products are actually sold rather than by running regressions, and no model will surface it.

Comparative disparate treatment asks whether similarly situated applicants were treated differently. This is where matched-pair analysis lives: find applicants whose credit characteristics are alike and whose outcomes differ, then read the files. It is labor-intensive, it is the most persuasive evidence in either direction, and statistics only tell you where to look.

Disparate impact asks whether a facially neutral policy produces a disproportionate adverse effect on a protected class. The long-standing analytical structure shifts burdens: a disparity is shown, the lender demonstrates a legitimate business necessity for the practice, and the question then becomes whether a less discriminatory alternative would serve that necessity. The legal status and enforcement posture of this theory have moved more than once in recent years, so treat the framework as an analytical tool and get current legal advice before treating it as a compliance obligation.

One practical note that survives any change in federal posture: state fair lending laws, state attorneys general, private litigation, investor representations and bank partner requirements do not move in lockstep with federal enforcement. A lender that stops testing because one enforcement channel narrowed has usually not narrowed its actual exposure.

You are usually not allowed to know

Here is the structural fact that shapes everything technical. Regulation B generally restricts a creditor from inquiring about an applicant's protected characteristics, with a specific exception requiring collection for certain dwelling-secured applications. That exception is why mortgage lenders have race and ethnicity data and why almost nobody else does.

So mortgage analysis is a data analysis problem, and everything else — auto, card, personal, small business — is an inference problem. Nobody chose this arrangement; it is the joint product of an anti-discrimination rule and a disclosure rule that were written for different purposes.

Where the data does not exist, the standard approach is proxy estimation. The best-documented method combines surname with geography using publicly available reference data to produce a probability distribution over race and ethnicity for each applicant, and it exists in the public record with a published methodology precisely so that lenders and supervisors can use the same technique. Its properties are worth stating honestly.

Proxy questionThe honest answer
Is it accurate for an individual?No. It produces a probability, not a fact, and it should never be attached to an individual decision or communicated as though it were known
Is it usable in aggregate?Generally yes, for portfolio-level disparity estimation, which is what the analysis needs
Does accuracy vary by group?Yes, substantially. Surname is far more informative for some groups than others, and geography carries most of the weight for the rest
Does geographic granularity matter?Yes. Finer geography carries more signal; coarser geography dilutes it toward the population average
Can it be used for sex or age?Those are usually available or inferable by other means; the proxy problem is mainly about race and ethnicity
Does using it create risk?Discuss it with counsel first. Estimating protected characteristics is a step with its own considerations, and the reason for doing it should be recorded

Two engineering consequences follow. Carry the probabilities through the analysis rather than assigning each applicant to their most likely group — hard assignment throws away information and biases estimates in ways that are hard to reason about. And propagate the proxy uncertainty into your confidence intervals, because a disparity estimated on probabilistic group membership is less precise than one estimated on reported data, and reporting it as though it were not is a quiet form of overstatement.

A proxy produces a probability, never a fact. It belongs in a portfolio-level estimate and never on an individual file, and the uncertainty it introduces belongs in every interval you report.

The controls that destroy the analysis

The mechanics are not exotic. For a binary decision, compare outcome rates across groups, then fit a model of the decision with legitimate credit factors included and look at what remains attributable to group membership. For pricing, do the same with price as the outcome. Report the unadjusted disparity and the adjusted one side by side, because the movement between them is itself the finding.

The judgment is entirely in the control variables, and this is where in-house analyses most often go wrong in a way that is invisible in the output.

Never control for a variable that is itself a channel of the conduct you are testing. If you are examining whether pricing discretion produced a disparity, controlling for the discretionary adjustment absorbs exactly the effect you are looking for and returns a reassuring zero. The same trap appears with loan officer identity, branch, channel and any downstream product assignment. These variables are worth analyzing as outcomes in their own right, and are usually wrong as controls.

Control for what the policy says drives the decision. Your written credit policy is the specification. If a factor is not in the policy but is in the model, that is a finding before any coefficient is estimated.

Segment before you conclude. A portfolio-level result of no disparity routinely conceals a real one in a single product, channel, region or period. Run the analysis by product, by origination channel, by geography and by time, and treat the segment view as primary rather than as a follow-up.

Watch the denominators. Withdrawn and incomplete applications are not denials, and how you treat them can move a result materially. Whatever you decide, write it down and apply it consistently across periods, because an unexplained change in that definition looks exactly like an attempt to move a number.

Where the findings tend to come from — our read

Discretionary pricing and exception practices
86
Geographic footprint, marketing and application flow
78
Inconsistent application of written credit policy
72
Adverse action reasons that do not match the model
60
No documented search for a less discriminatory alternative
56
A single model variable being the whole story
20

Our judgment of where issues concentrate, not a survey. Discretion and geography have produced findings for decades; the model is usually not the interesting part.

Statistical significance is not the standard

Two errors happen constantly and they run in opposite directions.

The first is treating a statistically significant result as automatically meaningful. With a large enough portfolio, a difference of a basis point or two will clear any significance threshold you like while meaning nothing to any applicant. Report the effect size in units a person can feel — percentage points of approval, dollars of price, basis points of rate — alongside the test statistic, and never present significance alone.

The second is treating a non-significant result as a clean bill of health. With four hundred applications, the analysis is nearly powerless and the honest statement is that the test could not detect a disparity smaller than some size. Compute and report that detectable size. It is the difference between “we found nothing” and “we could not have found anything,” and only one of those is a defensible sentence.

Two more points on measurement. Borrowing the four-fifths rule from employment selection guidance is common and is not a fair lending standard; if you use it as a screening heuristic, label it as one. And beware of the multiple comparisons you have made without counting them — running twenty segment tests will produce a significant result by chance, so pre-register the segments you intend to test, apply a correction, and treat the rest as exploratory pending confirmation.

The alternative search is the part nobody documents

If a disparity exists and the practice has a genuine business justification, the analysis is not over. The remaining question is whether something else would serve that justification about as well with less disparity. For a lender using models, this is a search you can actually run, and running it is what separates a program that is testing from one that is merely reporting.

The search has a definite shape. Fix the business objective — the performance the model must retain. Then generate genuine alternatives: a suspect variable removed or replaced, different functional forms, hyperparameters, regularization, training samples, and methods designed to reduce disparity subject to a performance constraint. Evaluate each on both axes and plot the frontier. If a candidate sits at materially lower disparity and near-identical performance, the interesting question becomes why you would not adopt it.

Two cautions. Techniques that adjust outputs by estimated protected class after the fact raise their own legal questions and should not be implemented on an engineer's judgment. And a variable-level review belongs alongside the model search: for each input, ask whether it has a plausible connection to creditworthiness and how strongly it correlates with protected class. Variables that are highly predictive of group membership and weakly connected to repayment are where scrutiny concentrates, and educational and geographic proxies have drawn attention for exactly that reason.

Whatever you find, document the search itself — the alternatives considered, how they were generated, the performance and disparity of each, and the reasoning for the selection. An undocumented search is indistinguishable from no search, and the record is the artifact a reviewer asks for.

Adverse action reasons have to be real

Denials require a statement of the specific principal reasons, and this obligation does not relax because the model is complicated. A reason drawn from a generic list because it seemed close is a defect on its own terms, and it also tells an applicant nothing they can act on, which is the point of the requirement.

Practically, reason codes must be derived from the actual decision: compute each input's contribution for that specific applicant, map contributions to plain-language reasons, and verify the mapping on real cases rather than in the abstract. Guidance on how this applies to complex models has been issued and revised in recent years, so confirm the current position — but the statutory requirement for specific, accurate reasons is the stable part, and it is what to design to.

Where a model cannot produce a defensible reason for an individual decision, that is a design constraint rather than a documentation problem, and it is one of the few places where model choice genuinely is dictated by the regulation.

Governance, and one thing worth knowing before you start

Run the analysis on a fixed cycle with a written plan: which products, which bases, which tests, which thresholds, and what happens when a threshold is crossed. Give it a named owner outside the business line, keep the results and the remediation record, and re-run after any material change to policy, model or footprint. Pre-implementation testing before a new model goes live is worth more than any amount of after-the-fact analysis.

One legal point engineers should know exists before running anything. Regulation B provides a privilege for certain lender self-tests, and its scope is narrower than most people assume — it is aimed at tests creating information not otherwise available in loan and application files, so a mystery-shopping exercise may qualify while a statistical analysis of data you already hold generally does not, and conditions are attached. Settle this with counsel before the work starts, because it may change how the analysis is scoped and who receives the output.

If you are small, do something different

If you originate a few hundred loans a year, portfolio statistics will tell you almost nothing, and building a fair lending analytics function is the wrong use of your money. This is a case where the honest advice is that you do not need an outside data firm.

What reduces risk at that size is consistency. A written credit policy specific enough that two officers reach the same decision. Exceptions logged with a reason and reviewed by someone who did not grant them. Pricing that follows a documented rate sheet, with discretion bounded and recorded. Complete application records, including the ones that went nowhere. And a periodic file review comparing similar applicants with different outcomes, done by hand on a sample — at your volume, more informative than any regression.

That is a compliance program, and it is what a reviewer would rather see from a small lender than a statistical report the institution cannot explain.

The mistakes we see

  • Controlling for the discretion being tested, which absorbs the disparity and returns a reassuring zero
  • Hard-assigning proxy groups instead of carrying probabilities through the estimate
  • Ignoring proxy uncertainty in the confidence intervals that get reported
  • Reporting significance without effect size, or effect size without power
  • Reading a non-significant result as an all-clear in a portfolio too small to detect anything
  • Testing at portfolio level only, where a real product- or channel-level disparity disappears
  • Running twenty segment tests and reporting the one that came out significant
  • No documented alternative search, which is indistinguishable from no search
  • Reason codes from a generic list rather than from the model's actual behavior
  • Changing a denominator definition between periods without recording why

Before the next testing cycle

  • Counsel has set the scope, the privilege position and who receives the output
  • A written testing plan names products, bases, tests and thresholds in advance
  • Proxy probabilities are carried through, never hard-assigned
  • Proxy uncertainty is propagated into reported intervals
  • Controls come from the written credit policy, and none is a channel of the conduct tested
  • Unadjusted and adjusted disparities are reported side by side
  • Results are segmented by product, channel, geography and period
  • Effect sizes are in units a person can feel, alongside the test statistic
  • Minimum detectable effect is reported wherever a result is null
  • An alternative search was run, and the alternatives and reasoning are documented
  • Reason codes are derived from the model and verified on real cases
  • Pre-implementation testing runs before any new model or policy goes live

Bottom line

Fair lending analysis fails for boring reasons. A control variable that absorbs the thing being measured. A null result from a sample too small to detect anything, reported as an all-clear. A proxy treated as a fact. An alternative search that happened in someone's head and was never written down. Fix those four and the analysis becomes useful to the business as well as to the examiner — because a disparity you find yourself is a design question with time to answer it, and one somebody else finds is a matter with a deadline attached.

Frequently asked questions

We do not collect race or ethnicity. Can we still test?

Yes, using proxy estimation from surname and geography, which is the standard approach outside mortgage and has a published methodology. The results support portfolio-level disparity estimates and should never touch an individual file or decision. Talk to counsel before you start: estimating protected characteristics is a deliberate step with its own considerations, and the reason for taking it belongs in the record.

Does using a machine-learned model create fair lending risk on its own?

Not by itself. The obligations are the same whichever method you use, and the disparities that draw findings usually come from discretion, geography and inconsistent policy rather than from the algorithm. What a flexible model changes is that it can find and exploit correlations you did not intend — which raises the value of variable-level review, a documented alternative search, and reason codes that reflect actual behavior.

How large a disparity matters?

There is no single number, and be wary of anyone who offers one. Materiality depends on the outcome, the size and direction of the effect, the population affected and the practice behind it. Report the effect in real units, report the confidence interval, report the sample and its power, and let a qualified reader judge. A well-presented small disparity with a plausible explanation is a very different document from a large one with none.

Should we just remove any variable that correlates with race?

Not mechanically. Nearly every useful credit variable correlates with protected class to some degree, and stripping them indiscriminately produces a weaker model without necessarily reducing disparity. The useful question is per-variable: how plausibly is this connected to repayment, how strongly does it predict group membership, and does removing it materially reduce disparity in a measured comparison? That is what the alternative search answers, and it answers it with evidence rather than instinct.

How often should this run?

On a stated cycle — many lenders test annually and monitor key indicators more frequently — plus event-driven testing whenever a model, policy, pricing structure or geographic footprint changes materially. The cycle itself should be in your written plan rather than decided each time. And do the pre-implementation test before a new model is live; it is the cheapest point at which a finding can still be a design decision.

1 business day response

Not sure whether your testing would hold up?

Send the testing plan and the variable list, and we will tell you plainly where the controls or the power look like a problem. Email bo@precisionfederal.com.

Email an engineerCapabilitiesMore insights →
Fair LendingDisparity TestingProxy MethodsAdverse Action