PAN Lab example
Apple Card underwriting
Cleared on the numbers but unable to say why
In 2019, people complained the Apple Card gave women lower credit limits. New York's regulator found no unlawful discrimination but faulted customer service and transparency.
See more
The Apple Card underwriting model is Goldman Sachs Bank USA's automated system for the credit card it offers with Apple. It decides each application and sets each approved applicant's credit limit and interest rate. Cardholders had little insight into why it set the terms it did.
How the complaints began
Apple and Goldman Sachs launched Apple Card in August 2019. Millions of Americans applied in the first few months, including hundreds of thousands of New Yorkers.
On November 7, 2019, a tech entrepreneur began tweeting that his Apple Card credit limit was 20 times his wife's, although they filed joint tax returns. He blamed sex discrimination. He also complained that a customer service agent could not explain the gap.
A co-founder of Apple tweeted that he, too, had been offered far better terms than his wife. Other consumers complained that women were offered lower limits and denied accounts unfairly. Within days, Goldman Sachs re-reviewed both wives' credit files and raised their limits to match their husbands'.
What the regulator examined
The New York State Department of Financial Services, the state's financial regulator, investigated the viral allegation. It reviewed several thousand pages of records and written answers from Goldman Sachs and Apple. It interviewed witnesses and complainants.
It also analyzed underwriting data for nearly 400,000 New York applicants. The data covered applications from the card's launch to the first complaints.
What the regulator found
Its report, published on March 23, 2021, found no unlawful discrimination on a prohibited basis, such as sex. Women and men with equivalent credit characteristics had similar outcomes. Goldman Sachs had a fair-lending program meant to keep the model from considering prohibited traits.
When the regulator asked, Goldman Sachs explained the decision for every consumer who had complained to the regulator. It named factors such as credit score, debt, income, credit utilization, and missed payments. None was an unlawful basis for a decision.
The regulator found that spouses who complained typically had different credit scores and credit histories. For example, one spouse was named on a mortgage and the other was not. Differences like these can lead to different credit offers.
What the regulator faulted
The same investigation found deficiencies in customer service and transparency. The regulator said deficiencies in customer service and a perceived lack of transparency undermined consumer trust in fair credit decisions.
The regulator noted that the law required a lender to explain its decision only when it denied credit, not the credit limit and terms it granted. Even so, the regulator said cardholders' confusion about their terms could have been reduced. Cardholders also had to wait six months to appeal their terms, until Goldman Sachs dropped the wait after the complaints.
In the rush to launch, the regulator wrote, Goldman Sachs seemed unprepared for a complex model producing outcomes that might surprise applicants. Internal deadlines and pressure to launch by a set date made this worse. The report says Goldman Sachs and Apple have since taken steps to remedy the deficiencies.
What changed
In June 2020 Goldman Sachs and Apple started Path to Apple Card, a program that gives declined applicants steps toward approval. By March 2021, more than 70,000 people had enrolled and nearly 5,000 had been approved. The Apple Card website had also begun to explain what data Goldman Sachs uses to set credit terms.
Two sides of one deployment
The case file splits the deployment into two sides that can fail on their own. One is the model and its fair-lending testing, which passed. The other is explaining decisions to the people they affect, which the regulator faulted.
Goldman Sachs could explain the decision for every consumer who had complained to the regulator. But a lack of transparency to the complainants themselves seemed to produce confusion, the regulator said. As the case file puts it, evidence that a model is fair is not evidence that the organization can explain its decisions.
The duty to explain
Federal credit rules require a lender that denies credit to give specific, accurate principal reasons, meaning the main reasons. The Consumer Financial Protection Bureau is a federal regulator of consumer credit. In guidance documents called circulars, from May 2022 and September 2023, it said this holds however complex the model is. A black box, a model too complex for its user to explain, is no defense. Checking the closest reason on the Bureau's sample form, a list of common denial reasons, does not comply.
These circulars came after the Apple Card report. The report does not say the card's denial notices fell short.
What the regulator concluded
The regulator wrote that credit scoring can reflect and carry forward past discrimination, even when it follows the law. It concluded that credit scoring and the laws against lending discrimination need strengthening and modernization to improve access to credit.
What this network is drawn from
This network follows the pattern the case file describes. It does not reconstruct Goldman Sachs's actual model. It shows the model, customer service, the credit policy and fair-lending team, and Goldman Sachs's decision and inquiry records. It also shows an explanation channel for applicants, and the adverse action notice, which a lender must send when it denies credit.
The network has two checks: the fair-lending review of the model, and a check on the explanations customer service gives. The sources describe no such explanation check at Apple Card. Applicants and cardholders are outside the network.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A tool closes a pathway when mistakes stop passing along it. The work along it may go on.
This case has a budget of 9 units and starts with no pressure applied. Explore (No Targets) sets no targets. The other three levels all ask for the network's mistakes to be contained, meaning corrected rather than building on each other.
Under Service Targets Only, the targets can be met with 85 different sets of tools, or 265 ways once stronger settings are counted. The cheapest costs 2 units: Mark AI-written records alone. Gate record entries alone meets them for 3 units. So do Peer sharing rules and Assign a challenger, each alone at its stronger setting.
With lingering effects turned off in the Dynamics menu, Peer sharing rules at its standard setting also meets them alone, for 2 units. Lingering effects are effects that stay after their cause is gone, and the Lab opens with them on.
Service and Safety Targets and All Governance Targets also ask you to close every failure pathway, among other targets. Seven pathways are open before any tool is used. The first four are Decisions brought to customer service, Appeals passed to credit policy, Customer complaints reported to the model's owners, and Customer questions recorded. The other three are Fair-lending results recorded, Past decisions used to maintain the model, and Decision history read by customer service.
No tool on offer closes Customer complaints reported to the model's owners. So those two levels are not fully addressable with the available tools, even with the budget lifted.
More is not better here. Using every tool at its strongest setting costs 25 units, well over the budget of 9. It contains the mistakes and closes six of the seven failure pathways. But the benefit the model brings to the work falls below what the targets ask, so it meets them at none of the three levels.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Cleared-underwriting-class with the explanation channel unbuilt network: 6 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 5 assumptions
- assumed
The network draws the adverse action notice as its own part, between the model's decision and the applicant. The case turns on explaining decisions, so the notice is not folded into another part. Whether a notice's reasons are specific and accurate is a separate question from whether the model is fair. The regulator's report does not say the card's denial notices fell short. The network also assumes a heavy workload for limited staff, for a card millions applied for in its first months.
- baseline
This network follows the pattern the case file describes. It does not reconstruct the actual model. New York's financial regulator analyzed underwriting data for nearly 400,000 New York applicants and found no unlawful discrimination on a prohibited basis. The network draws that finding as the fair-lending review of the model, the deployment's genuine strength on the public record.
- baseline
The network draws the failure as a separate check on customer service, the explanation check. The regulator found deficiencies in customer service and transparency, even though the underwriting was lawful. It said deficiencies in customer service and a perceived lack of transparency undermined consumer trust. One complainant said an agent could not explain the gap between his and his wife's credit limits. Passing the fair-lending review did not give customers explanations.
- assumed
Federal credit rules require a lender that denies credit to give specific, accurate principal reasons, however complex its model. The Consumer Financial Protection Bureau has said a black box is no defense. It has also said that checking the closest reason on its sample form of denial reasons does not comply. The network draws the gap as the link between customer service and the credit policy team. Each holds part of an explanation the other would have to complete.
- assumed
The network shows no credit decision and no applicant. It shows how mistakes pass between the organization's parts. The finding of no unlawful discrimination, the documented transparency failures, and the loss of trust come from the case file, and nothing in the network computes them. The viral complaints that prompted the investigation are recorded as what started it, never as a verdict the network decides.
What this example does not show
Show all 2 limitations
- This example shows no credit decision and no applicant. It shows only how mistakes pass between the organization's parts. The regulator's finding of no unlawful discrimination, the documented transparency failures, and the loss of trust come from the case file and are not computed here.
- The regulator found the underwriting lawful but found deficiencies in customer service and transparency. The network draws these as two separate parts, a fair-lending review and an explanation check, not as a computed harm. The viral complaints that started the investigation are recorded as its trigger, never as a verdict.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A bank's automated credit-decisioning for a widely used consumer card was investigated by a state regulator after viral allegations of gender bias in credit-line assignment. The regulator analyzed roughly 400,000 in-state applicants and found no unlawful discrimination on a prohibited basis — the model was cleared on the numbers. But the same investigation documented failures of explanation, customer service, and perceived transparency: applicants had little insight into why they received the terms they did, a complainant said customer service could not explain the decisions, and the resulting opacity undermined consumer trust even though the underwriting itself was found lawful. This is the domain's cleared-but-faulted case: a statistically clean model paired with a failure of transparency.
empirical- Government evaluation New York State Department of Financial Services (2021, March 23). Report on Apple Card Investigation. https://www.dfs.ny.gov/reports_and_publications/press_releases/pr202103231
The lesson the cleared-but-faulted outcome carries is that a lawful, statistically clean model does not discharge the separate duty to explain a decision. Regulators have made explicit that adverse-action notices must give specific, accurate principal reasons regardless of how complex the model is, and that a model being a black box is not a defense — checking the nearest sample-form box does not comply. The explanation and customer-service channel is therefore a distinct, separately-resourced surface that can fail on its own: an organization can pass its fair-lending testing and still fail the people it decides on by not telling them why.
empirical- Regulatory Consumer Financial Protection Bureau (2022, 2023). Circular 2022-03: Adverse action notification requirements in connection with credit decisions based on complex algorithms; and Circular 2023-03 on Regulation B sample forms. https://www.consumerfinance.gov/compliance/circulars/circular-2023-03-adverse-action-notification-requirements-and-the-proper-use-of-the-cfpbs-sample-forms-provided-in-regulation-b/
Where this connects
Institutional pressures in this domain
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Lending & credit collections AI domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Gate record entries — Human-in-the-loop write gating
- Peer sharing rules — Peer-edge governance
- Review on schedule — Oversight cadence & retrospectives
- Assign a challenger — Structured dissent
- Escalate checks — State-feedback vigilance
- Upgrade model — Improve the model
Documented case histories
- Cleared on the numbers but faulted on the explanation
- Automated underwriting with its fair-lending testing on the record
- The governance an enforcement action had to write
- M-Shwari & Kenya's Digital Credit Market
- Citi Retail Services Judgmental Review & the Armenian surname screen
- Santander Consumer USA subprime vehicle loan scoring
- Credit Acceptance Corporation's net-collections score
- Wells Fargo refinance underwriting & the bridge nobody could build
- Navy Federal mortgage underwriting & three readings of one gap
- Enova International servicing defects & the debits nobody authorised
- Equifax Online Model Server coding error (2022)
- TransUnion's OFAC Name Screen & the people who could not sue
- Dave ExtraCash: an advertised ceiling, an automated amount, and a case that never asks how the amount is set
- Hello Digit's automated-savings algorithm
- Oportun's legal-collections filing pipeline