PAN Lab example
Upstart lending model
Regulator-verified access — and a search left at an impasse
Upstart's model approves, declines, and prices personal loans with no human review. Published reports show wider access to credit, and approval gaps for Black applicants.
See more
Upstart Network is a US lending platform that works with banks. Its machine-learning model decides whether to approve a personal loan and sets its price. It uses education and other alternative data, meaning information beyond the usual credit score, and no person reviews each application.
The no-action letter and its results
In 2017 the Consumer Financial Protection Bureau (CFPB), the US federal consumer-finance regulator, issued its first no-action letter, to Upstart. Such a letter tells a company that the regulator's staff has no present intention to recommend enforcement or supervisory action on a specific matter. Upstart's letter required it to report to the CFPB, and it ran until 2022.
In August 2019 the CFPB published highlights of Upstart's own simulations and analyses under the letter. The CFPB noted it had not separately replicated them. They compared Upstart's model with a traditional lending model.
Upstart's model approved 27 percent more applicants, at 16 percent lower average APRs for approved loans. The APR, or annual percentage rate, is the yearly cost of borrowing. Near-prime applicants, with FICO credit scores of about 620 to 660, were approved about twice as often.
Across the race, ethnicity, and sex groups tested, approvals rose by 23 to 29 percent and average APRs fell by 15 to 17 percent. Among its lending cases, the case file counts these as the only results on how well a lending system serves people that a regulator published.
The monitorship
Under a separate private agreement, Upstart, the NAACP Legal Defense Fund, and the Student Borrower Protection Center set up a fair-lending monitorship. A monitorship is an independent outside review, here of the live model. The monitor was the law firm Relman Colfax. It engaged Sentrana, a firm in machine learning and artificial intelligence, as a consultant. It also engaged Dr. Bernard Siskin of BLDS, an expert on statistical analyses that measure discrimination in financial services.
It published four reports, in April 2021, November 2021, September 2022, and March 2024. Its statistical tests found no close proxies for protected traits among the model's inputs. A proxy is an input that stands in for race or another trait that fair-lending law protects.
It did identify approval disparities for Black applicants. Its September 2022 report identified what would likely have been a viable less-discriminatory alternative. That is a different model with smaller gaps between groups, and this one appeared to perform comparably. Upstart updated its model before that analysis was finished.
The final report, in March 2024, recorded an impasse. Upstart, the NAACP Legal Defense Fund, and the Student Borrower Protection Center disagreed over the right, legally required way to judge comparable performance.
What a finding about inputs does not settle
Finding no close proxies is a finding about the inputs. It does not clear the outcomes. A feature that looks neutral can still affect protected groups unequally, even when no single input is a close proxy.
One example is a school's cohort default rate, the share of its former students who default, used to price one person's loan. Disparate-impact testing exists to catch this. It compares outcomes, such as approval rates, across the groups fair-lending law protects. The sources here do not say whether Upstart's model used that feature.
Who decides
Underwriting, deciding whether to approve a loan and at what price, is fully automated. So no person reviews a single application. The compliance and model-risk team tests the model's decisions in aggregate, across many applications.
Saying the model decided does not end accountability. It moves all of it to choices made before any decision: which model, what testing, how hard to look for alternatives, and what to report. The question is whether those choices are paid for and staffed, or only declared.
The duty to explain a denial
When the model declines someone, Upstart must still give that applicant specific, accurate principal reasons, meaning the main reasons for the denial. The duty comes from the Equal Credit Opportunity Act and its rule, Regulation B.
A circular is a CFPB guidance document. In Circular 2022-03, in May 2022, the CFPB said a model's complexity is no excuse. If a system cannot produce accurate reasons, the lender may not use it. Circular 2023-03, in September 2023, said its sample denial forms, with reasons to tick, do not by themselves comply.
This duty is separate from fairness testing. A lender can pass its bias tests and still fail here, by being unable to tell a declined applicant why.
What this network is drawn from
This network follows the pattern in the case file. It is not a reconstruction of Upstart's actual model. It shows the model, the compliance and model-risk team, the decision record, the monitorship, and the CFPB's reporting under the letter.
The applicants are outside the network, and no credit outcome is computed on it. Repayment results exist only for funded loans, so a wrong decline leaves no trace in them.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A tool closes a pathway when mistakes stop passing along it. The work along it goes on.
This case has a budget of 9 units. Understand the system, a tool that lowers other tools' prices, is not offered here, so every price is full. Every tool costs the same at every target level.
The case starts with one pressure on, Monitoring goes stale. Both scheduled outside looks at this model have ended: the no-action letter in 2022, and the monitorship with its final report in March 2024.
The network starts at a tipping point, the Lab's term for a network where mistakes could either die out or build on each other. The aim is self-correcting: mistakes are corrected rather than building on each other.
Explore (No Targets) sets no targets. There, no single tool makes the network self-correcting. Pairs do, and the cheapest is Review on schedule with Assign a challenger, for 4 units.
Under Service Targets Only, the targets can be met. That level asks for the network to be self-correcting, and for the model to be helping the work enough. The model already helps enough at the start, so the mistakes are what must change.
The Lab starts with its Dynamics setting at Both, where side effects and lingering effects are both on. At that setting, 39 combinations of tools and settings meet these targets within the budget. They use 13 different sets of tools.
The cheapest is the same pair, for 4 units. Assign a challenger with Store less data, or with Upgrade model, costs 5 units. No single tool meets the targets alone.
With Dynamics at Off, 49 combinations meet them, using 20 sets of tools. There, Assign a challenger with Check with a second model also costs 5 units, and so does Review on schedule with Store less data.
Service and Safety Targets and All Governance Targets are not fully addressable with the available tools. Both levels ask you to close every failure pathway, among other targets. Seven are open at the start.
Four are "Decision results for testing", "Testing findings shape the model", "Test results filed", and "Repayment data retrains the model". The other three are "Decision record read for testing", "Monitorship reads testing evidence", and "Results reported to the CFPB".
Store less data closes "Test results filed". No tool offered here acts on the other six. Every combination was checked, within the budget and with the budget lifted, and none meets these two levels.
That is a finding about the deployment, not a gap in your approach. Those links are how the team, the record, the model, and the two outside reviewers do their work.
More is not better here. Using every tool at once, each at its stronger setting, costs 22 units, well over the budget of 9. It makes the network self-correcting. But the model's benefit to the work falls below what the targets ask, so it meets them at no level.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Automated-underwriting-class with its fair-lending testing on the record network: 5 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 7 assumptions
- assumed
The monitorship and the CFPB's reporting are shown as two separate outside reviewers, because the sources describe two with different authority. The monitorship came from a private agreement with the NAACP Legal Defense Fund and the Student Borrower Protection Center. It examined the live model and published four reports. The reporting came from the CFPB's no-action letter and ran from 2017 to 2022. The CFPB published figures from it. Merging the two would hide that they watched the same model from different places. The network also assumes more work than the staff can handle. The sources do not give the team's size. On a national lending platform, the team tests in aggregate and never sees an application.
- baseline
This network follows the lending pattern in the case file. It is not a reconstruction of the actual model. Under the no-action letter, Upstart had to report to the CFPB. In August 2019 the CFPB published highlights of Upstart's own simulations, which it did not separately replicate. Against a traditional model, they showed 27 percent more applicants approved at 16 percent lower average APRs. Near-prime applicants were approved about twice as often, with gains across the groups tested. The network treats these as Upstart's figures that a regulator published, not as figures a regulator checked.
- assumed
Underwriting, deciding whether to approve a loan and at what price, is fully automated, with no person reviewing each application. So the network's human part is the compliance and model-risk team, which tests the model across many decisions. Saying the model decided does not end accountability. It moves it to choices made upstream: the model, the testing, the search for alternatives, and the reporting. The question is whether those choices are paid for and staffed, or only declared.
- baseline
The network includes the fair-lending testing because four public monitorship reports document it on the live model. The monitorship's statistical tests found no close proxies for protected traits among the inputs. It also identified approval disparities for Black applicants. So a finding of no proxies is about inputs, and it does not clear the outcomes. A neutral-looking feature can still affect protected groups unequally. One example is a school's cohort default rate, the share of its former students who default, used in one person's price. Disparate-impact testing exists to catch that.
- assumed
The network includes the monitorship's review of an alternative model as a check. The monitorship identified what would likely have been a viable less-discriminatory alternative that appeared to perform comparably. Upstart updated its model before that analysis was finished. The final report recorded an impasse over the legally required way to judge whether such an alternative performs comparably. The sources show that question left open, not settled.
- baseline
The duty to explain a denial is placed on the pathway where decisions and denial reasons are logged. When the model declines an applicant, Upstart must still give specific, accurate principal reasons. The CFPB has said a model's complexity is no excuse. A lender can pass its bias testing and still fail this duty, by being unable to tell a declined applicant why. So the network treats explaining as its own duty, not a side effect of fairness testing.
- assumed
The network models no credit outcome and no applicant. It shows how mistakes move inside the organization only, and applicants are outside it. Approvals, declines, disparity findings, the impasse over alternatives, and the reasons on any denial notice are in the case file. None is computed from this network.
What this example does not show
Show all 2 limitations
- This example models no credit outcome. It shows how mistakes move inside the organization only, and applicants are outside the network. The access figures, the monitorship's findings, the impasse over alternatives, and the reasons on any denial notice are in the case file. None is computed from this network.
- The 27 percent and 16 percent access figures come from Upstart's own simulations, which the CFPB published in 2019 without replicating them. The no-proxies finding, the approval disparities, and the impasse over alternatives come from the monitorship's public reports. The duty to explain a denial is shown on the pathway where decisions are logged, not as a computed harm.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A machine-learning underwriting and pricing platform using education and other alternative data operated for five years under a regulator's no-action letter with a reporting obligation, and the regulator published the access results: 27 percent more applicants approved than a traditional model at 16 percent lower average APRs, with near-prime applicants (FICO 620 to 660) approved at roughly twice the rate, and gains across the tested demographic segments. This is the lending family's only regulator-published service term, drawn from the company's own simulations, which the regulator did not separately replicate. Underwriting is fully automated with no per-application human review, so the organizational levers are all upstream — model choice, the testing regime, the search for alternatives, and the reporting channel to the regulator.
empirical- Government Consumer Financial Protection Bureau — Ficklin, P.A., & Watkins, P. (2019). An update on credit access and the Bureau's first No-Action Letter. CFPB Blog. https://www.consumerfinance.gov/about-us/blog/update-credit-access-and-no-action-letter/
The same deployment carries the family's most detailed public fair-lending testing record: four reports from an independent monitorship agreed with civil-rights organizations found no close protected-class proxies quantitatively, but identified approval disparities for Black applicants, flagged a likely viable less-discriminatory alternative model, and ended in a documented methodological impasse over the legally required way to judge whether such an alternative performs comparably. Independently of the disparity question, adverse-action notices must give specific, accurate principal reasons for a denial regardless of the model's complexity — a governed explanation duty a complex model does not discharge by being accurate.
empirical- Advocacy Relman Colfax PLLC (2021-2024). Fair Lending Monitorship of Upstart Network's Lending Model (Initial, Second, Third, and Final Reports). https://www.relmanlaw.com/cases-upstart-network-fair-lending-counseling
- Regulatory Consumer Financial Protection Bureau (2022, 2023). Circular 2022-03: Adverse action notification requirements in connection with credit decisions based on complex algorithms; and Circular 2023-03 on Regulation B sample forms. https://www.consumerfinance.gov/compliance/circulars/circular-2023-03-adverse-action-notification-requirements-and-the-proper-use-of-the-cfpbs-sample-forms-provided-in-regulation-b/
Where this connects
Institutional pressures in this domain
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Lending & credit collections AI domain page.
Levers available here and the patterns behind them
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Assign a challenger — Structured dissent
- Store less data — Data minimization
- Upgrade model — Improve the model
Documented case histories
- Automated underwriting with its fair-lending testing on the record
- Cleared on the numbers but faulted on the explanation
- The governance an enforcement action had to write
- M-Shwari & Kenya's Digital Credit Market
- Citi Retail Services Judgmental Review & the Armenian surname screen
- Santander Consumer USA subprime vehicle loan scoring
- Credit Acceptance Corporation's net-collections score
- Wells Fargo refinance underwriting & the bridge nobody could build
- Navy Federal mortgage underwriting & three readings of one gap
- Enova International servicing defects & the debits nobody authorised
- Equifax Online Model Server coding error (2022)
- TransUnion's OFAC Name Screen & the people who could not sue
- Dave ExtraCash: an advertised ceiling, an automated amount, and a case that never asks how the amount is set
- Hello Digit's automated-savings algorithm
- Oportun's legal-collections filing pipeline