Most fair lending programs were built to test scoring models and underwriting algorithms. They were not built to test a chatbot having a conversation.
Generative AI is showing up in lending faster than compliance programs have adapted to it, usually through customer facing tools, AI underwriting assistants, and automated communication rather than the underwriting model itself. That matters because generative artificial intelligence creates a kind of fair lending risk that looks nothing like the statistical bias most testing programs were designed to catch. A chatbot does not produce a denial rate. It produces language, and that language can vary from consumer to consumer in ways no one designed and no one is testing for.
Here is where these tools are actually turning up, why the risk is different, and what the rules require now that a lot of the federal guidance has been pulled back.
Jump to section
- Key Takeaways
- Where Generative AI Is Actually Showing Up in Lending
- Why Generative AI Creates a Different Kind of Risk
- Specific Blind Spots to Watch For
- Why Existing Fair Lending Controls Do Not Cover This
- Where the Rules Actually Stand Right Now
- What Financial Organizations Should Do Now
- Frequently Asked Questions
- The Bottom Line
- How RADD Can Help
- Start With an Inventory
Key Takeaways
- Generative AI usually enters lending through chatbots, AI underwriting assistants, adverse action language, and marketing, not through the credit model itself.
- The risk shows up as inconsistency rather than statistical disparity, which is why outcome based fair lending testing misses it.
- The CFPB withdrew its algorithmic adverse action circulars on May 12, 2025. ECOA and Regulation B did not change.
- Enforcement pressure has shifted to state attorneys general, prudential examiners, and private plaintiffs.
- Revised federal model risk guidance issued April 17, 2026 explicitly excludes generative and agentic AI from its scope.
- The practical fix is an inventory, consistency testing, structured logging, and a review checkpoint before deployment.
Where Generative AI Is Actually Showing Up in Lending
Customer facing chatbots are the most visible example, handling loan inquiries, application status questions, and basic eligibility screening. Many consumers now interact with a generative tool well before they ever reach a human, sometimes without realizing it.
AI underwriting assistants sit further inside the process, summarizing applications, drafting explanations for human underwriters, or helping staff interpret model output faster. These tools do not make the credit decision themselves, but they shape what underwriters see and how they see it.
Generative tools are also being used to produce adverse action explanations, translating structured model output into consumer facing language. That use case sits directly at the intersection of explainability and fair lending, because the accuracy of the generated language carries real regulatory weight.
Marketing and outreach is the fourth area, with generative tools drafting or personalizing messaging sent to prospective borrowers.
What ties these together is that adoption is fastest in customer facing and support functions, which historically sat outside the model governance process. That is exactly why oversight lags. A tool bought to reduce call center volume is rarely the first thing anyone thinks to put through fair lending testing.

Why Generative AI Creates a Different Kind of Risk
Traditional model risk is about a model producing systematically different outcomes for different groups. That is the world of credit scorecards and machine learning models, and the risk is measurable in familiar terms: approval rates, pricing, scoring differentials.
Generative AI risk looks different. It shows up as inconsistency rather than statistical disparity. The same chatbot gives different quality explanations, different levels of detail, or a different tone to different consumers asking substantially the same question, with no policy driving the variation.
Generative tools also do not follow a fixed decision tree. Large language models are probabilistic by design, so the same prompt can produce different phrasing each time, which makes the output harder to test, harder to document, and harder to defend if a regulator asks why a specific consumer received a specific response.
There is also a blurring effect between decisioning and communication. When a generative tool explains an adverse action or answers a credit related question, the line between customer service and credit decisioning gets hard to locate. A chatbot that subtly discourages certain applicants from continuing an application is functioning as part of the credit decision whether or not anyone designed it that way.
Training data adds one more layer. Tools trained or tuned on internal communications, past customer interactions, or general internet text can absorb and reproduce subtle language patterns, including differences in tone or formality that track how certain groups have historically been addressed. None of that has to be intentional to create exposure.
The two risks compared:
| Traditional model bias | Generative AI risk | |
|---|---|---|
| What it looks like | Statistical disparity in outcomes | Inconsistency in language and detail |
| Where it lives | Approval rates, pricing, scores | Chat transcripts, summaries, notices |
| How you find it | Disparate impact analysis | Consistency testing across phrasings |
| Who deployed the tool | Model development | Customer service, marketing, ops |
| Audit trail | Structured and reviewable | Often missing entirely |
Specific Blind Spots to Watch For
Chatbots giving inconsistent guidance based on how a question is phrased. Two consumers asking the same underlying question in different language, or with different levels of financial fluency, can get meaningfully different answers from the same tool.
Underwriting assistants introducing bias through summarization. A tool that summarizes an application for underwriters can emphasize or omit details in ways that shape judgment, usually without leaving a record of what was included or left out.
Generated adverse action language that varies in specificity. If a generative tool is writing denial explanations, the clarity of those explanations can vary from consumer to consumer. That is a UDAAP transparency issue and a fair lending documentation issue at the same time, especially if the variation tracks any group characteristic.
Steering through tone and framing. A chatbot that frames a product more or less favorably based on patterns absorbed from training data creates exposure that numeric disparate impact testing will never catch, because there is no number to test.
No clear audit trail. Many generative tools do not produce the structured, reviewable output that underwriting systems generate by default. That makes it hard to reconstruct why a particular consumer got a particular response, which becomes a problem the moment a complaint or an exam asks that exact question.

That last pattern is the reason consumer complaints often surface AI fair lending problems before any testing program does. When the tool leaves no trail, the consumer becomes the record.
Why Existing Fair Lending Controls Do Not Cover This
Most fair lending testing is built around structured, numeric model output. Generative tools produce natural language, and most testing infrastructure has no built in way to evaluate that output for consistency or fairness.
Model risk frameworks built for statistical and machine learning models often have no category for large language models at all, especially ones adopted quickly by customer facing or marketing teams rather than by a model development function. A framework designed to validate a credit scoring model does not know what to do with a chatbot.
These tools are also deployed by teams outside the normal compliance review process. A tool bought to handle customer inquiries or draft marketing copy reads as a productivity tool, not a credit decisioning tool, so it can go live without ever passing the review that would catch a traditional underwriting model.
Vendor provided tools add opacity on top of that. A financial organization licensing a third party chatbot may have limited visibility into how the underlying model was trained or tuned, which makes consistency risk hard to assess even when the organization wants to.
Where the Rules Actually Stand Right Now
This is the part most coverage of the topic gets wrong, and it cuts both ways.
What the CFPB withdrew, and what it did not
In May 2022 the CFPB issued Circular 2022-03, stating that lenders using complex algorithms remain fully responsible under ECOA and Regulation B for specific, accurate adverse action reasons. It repeated the point in Circular 2023-03. The Bureau also published an issue spotlight on chatbots in consumer finance in June 2023, flagging accuracy and consumer harm risks in automated customer service.
Both circulars were withdrawn on May 12, 2025, as part of a withdrawal of dozens of guidance documents. The Bureau has not said the underlying obligations changed. Its April 2025 priorities memo puts mortgages first, followed by consumer reporting, debt collection, and fraudulent fees, which tells you where examiner attention is likely to go, not what the law now permits.
What did not change is the law. Regulation B still requires specific reasons for adverse action at 12 CFR 1002.9(b)(2), which rules out generic explanations and internal scoring cutoffs. ECOA still reaches every aspect of a credit transaction, including how a lender communicates with applicants. The Fair Housing Act still applies to mortgage lending regardless of who or what drafted the language. Withdrawing an interpretation of a rule does not withdraw the rule.
Who enforces this now
The practical effect of the withdrawal is not less risk. It is different risk, spread across more parties.
State attorneys general are not bound by the Bureau’s enforcement priorities, and several have been active on algorithmic lending. Prudential examiners at the banking agencies and the NCUA still review compliance management systems, and a generative tool sitting outside your governance framework is an obvious finding. Private plaintiffs can bring ECOA and Fair Housing Act claims directly. None of that depends on a CFPB circular being in force.
For a compliance officer, the practical read is simple. The federal guidance that told you what good looked like is gone. The obligation it was interpreting is not. You now have less written cover for your approach and the same underlying exposure, which is a worse position, not a better one.
What changed in 2026
Two developments since the withdrawal matter more than the withdrawal itself.
On April 17, 2026 the OCC, Federal Reserve, and FDIC jointly issued revised model risk management guidance, replacing SR 11-7, OCC Bulletin 2011-12, and FDIC FIL-22-2017. The revision explicitly excludes generative and agentic AI from its scope, describing those technologies as novel and rapidly evolving, and directs institutions to rely on broader risk management and governance practices instead. Read that carefully. The framework a compliance officer would reach for to govern a chatbot now says it does not cover chatbots. The agencies have said they intend to issue a request for information on banks and their use of AI.
Colorado moved the other way. Senate Bill 26-189 was signed on May 14, 2026 and is set to take effect January 1, 2027, though that date is subject to pending litigation. For consequential decisions involving lending, covered entities have to give clear and conspicuous notice before an automated system is used, and within 30 days of an adverse outcome provide a plain language explanation of the decision and the role the system played in it. Consumers can request access to and correction of inaccurate personal data used in the decision, and meaningful human review where commercially reasonable.
The combination is what to plan around. Federal model risk guidance now carves generative AI out, and at least one state is about to require a plain language explanation of it on a thirty day clock.
What Financial Organizations Should Do Now
Bring generative AI tools into existing model governance. Any tool touching customer communication, underwriting support, or adverse action explanation should be inventoried and reviewed the same way a credit scoring or machine learning model would be, regardless of which team deployed it.
Test for consistency, not just statistical disparity. Build protocols that probe whether a chatbot or assistant gives consistent quality and content across different types of consumers asking substantially similar questions, rather than relying only on outcome based metrics.
Require structured logging of generative outputs. Make sure these tools produce a reviewable record of what was said to whom, so patterns can be tested later and specific interactions can be reconstructed when a complaint or exam requires it.
Set boundaries on what generative tools can and cannot do. Draw a clear line between customer service and credit guidance, and build guardrails that stop a tool from drifting into de facto decisioning.
Scrutinize vendor tools as closely as internally built ones. Apply the same fair lending due diligence you would apply to a purchased underwriting model, including direct questions to vendors about testing methodology, training data, and consistency controls.
Bring legal and compliance in before launch, not after. Given how fast these tools get adopted by customer facing teams, build a review checkpoint before deployment rather than relying on compliance to catch problems once the tool is already talking to consumers.
Frequently Asked Questions
Does fair lending law apply to chatbots and generative AI tools used in lending?
Yes. Fair lending laws apply to the full lifecycle of a credit transaction regardless of whether artificial intelligence is involved, including customer service and communication, not just the underwriting decision. A chatbot that discourages certain applicants, gives inconsistent guidance, or shapes how consumers understand their options can create exposure even though it never calculates a credit score.
How is generative AI fair lending risk different from traditional model bias?
Traditional model bias shows up as statistical disparity in structured outcomes such as approval rates or pricing. Generative AI risk shows up as inconsistency in unstructured output, such as language, tone, or level of detail varying from consumer to consumer. Most existing testing was built to catch the first kind, not the second.
Did the CFPB withdrawing its AI circulars remove the adverse action obligation?
No. Circulars 2022-03 and 2023-03 were withdrawn on May 12, 2025, but they were interpretations of an existing rule. Regulation B at 12 CFR 1002.9(b)(2) still requires specific, accurate reasons for adverse action, and ECOA still applies. Withdrawing an interpretation does not withdraw the rule. What changes is where enforcement attention is likely to come from: state regulators, prudential examiners, and private litigation rather than the Bureau.
Can a chatbot create fair lending risk even without making credit decisions?
Yes. A tool does not need to approve or deny credit to create exposure. If a chatbot’s responses steer certain consumers away from applying, or give some consumers less complete information than others, it is functioning as part of the credit process in practice.
What should financial organizations test for when reviewing generative AI tools?
Consistency is the key concept. Test whether the tool gives comparable quality, completeness, and tone across different phrasings, different levels of financial fluency, and different consumer profiles asking substantially similar questions, on top of any accuracy testing already in place.
Do adverse action notice requirements apply to AI generated explanations?
Yes. Lenders remain responsible for providing specific, accurate reasons for adverse action even when a generative tool drafts the explanation. A model that is proprietary or hard to interpret does not relieve a lender of that requirement, and neither does the fact that software wrote the sentence.
Who inside a financial organization should review generative AI tools before deployment?
Compliance and legal should be involved before any generative tool touching customer communication or lending support goes live. Because these tools are often adopted by customer facing or marketing teams outside the model governance process, organizations need a defined checkpoint so fair lending review happens regardless of which department introduces the tool.
The Bottom Line
Generative AI has introduced a genuinely new category of fair lending risk, built around inconsistency and unstructured output rather than statistical disparity. A chatbot does not produce a denial rate a testing program can flag. It produces language, and language can vary from consumer to consumer in ways just as capable of creating discriminatory outcomes as any scoring model.
Most fair lending programs were not built with this in mind, and most generative tools were not deployed with fair lending review in mind either. That combination is exactly how a blind spot forms: a tool nobody classified as a credit decisioning system, reviewed by a program that was never built to test conversational output.
The federal guidance that used to spell this out has been withdrawn. The statutes it interpreted have not. Organizations that extend governance and testing to cover generative AI now will be ahead of both the regulatory curve and the risk itself.
How RADD Can Help
RADD helps financial organizations close this gap before it becomes a finding.
Generative AI inventory and risk assessment. We identify every generative tool touching customer communication, underwriting support, or adverse action explanation, including tools deployed outside the traditional model governance process, and assess each one for fair lending exposure.
Consistency testing frameworks. We design testing protocols built specifically for unstructured, natural language output, evaluating whether a chatbot or assistant delivers consistent quality, completeness, and tone across different consumers asking substantially similar questions.
Adverse action language review. We review AI generated denial explanations for specificity and accuracy, so they meet the Regulation B standard regardless of whether a human or a generative tool drafted the language.
Vendor due diligence support. We help evaluate third party generative AI tools with the same rigor applied to a purchased underwriting model, including targeted questions for vendors about training data, tuning, and consistency controls.
Start With an Inventory
You do not need a new testing program to begin. You need a list of every generative tool that talks to a consumer or shapes a credit decision, and an honest answer about which ones have never been reviewed.
RADD’s compliance consulting and fintech compliance services start exactly there. We build the inventory, test the tools against fair lending standards, and give you a written record your board and your examiners can follow.
Get a quote and tell us what generative tools you are running. We will tell you which ones your fair lending program is not covering.
