Skip to content

PAN Lab example

NYC MyCity business chatbot

Exposure is not correction: a public-facing government advice chatbot

New York City's MyCity chatbot said business owners could break laws protecting workers and tenants. It stayed online roughly two years after reporters exposed it.

See more

The MyCity Business chatbot was a generative AI chatbot that New York City launched in 2023 on Microsoft Azure AI. The city's Office of Technology and Innovation ran it. It answered business owners' questions about city rules, and it was confidently and repeatably wrong about what the law required.

What the reporting found

On 29 March 2024, The Markup published an investigation with THE CITY and Documented NY. It found the chatbot routinely advised businesses in ways that would break the law.

It said employers could take a cut of workers' tips. It said landlords could refuse tenants who pay with Section 8 housing vouchers or other sources of income. New York City law forbids that. It said tenant lockouts were allowed, and that there were "no restrictions" on rent charges. It said stores could go cashless, against a 2020 city law. It said funeral homes could conceal their prices, against the federal funeral rule.

The reporting tested specific questions. It did not measure an error rate across all answers.

When ten staffers at The Markup asked the housing-voucher question, all ten got the same wrong answer. An earlier test had returned the correct answer, so the answer had changed over time.

How the city responded

Mayor Eric Adams declined to take the chatbot offline and called it a pilot. "It's wrong in some areas, and we've got to fix it," he said. He also said: "we're going to have the best chatbot system on the globe."

The city relabeled the chatbot a beta product and added a disclaimer. The Office of Technology and Innovation said the chatbot had "already provided thousands of people with timely, accurate answers." That claim was not independently confirmed. Microsoft said it was working with the city on fixes.

About two weeks after the tips error came to light, the city applied a patch. It sent questions outside the chatbot's topics to NYC.gov instead of answering them.

In March 2025, then chief technology officer Matt Fraser said the city still planned to expand the chatbot. It would cover all the content of the city's 311 service, not small business alone. He said a patch had cut hallucinations, meaning made-up answers, "exponentially." He gave no figure.

What the audit found

On 30 December 2025, the Office of the New York City Comptroller issued a performance audit of the MyCity system. Brad Lander was then the city's Comptroller. The audit reviewed the system's cost, its planning, and whether it did what was promised, including the chatbot's answers. It found the chatbot "appears to be unable to provide accurate or consistent information."

An internal weekly production report, reproduced in the audit, showed the chatbot did not answer 23 of 48 tested government questions. People asked more than 2,200 questions in July and August 2025. Of the 70 users who left thumbs-up or thumbs-down feedback, 71.4 percent were negative (50 of 70). The city disputes that share, putting it at roughly 2.25 percent of all responses. The audit rebuts the city's way of counting.

The auditors' own testing found the same question answered two different ways. Answers also changed with trivial edits, such as "NYC" versus "nyc."

The audit also found the wider MyCity system had cost over $100 million, across more than 120 agreements with about 50 vendors. It lacked a system development plan. It had not delivered the promised single form for reaching city benefits.

The audit made seven recommendations, including one to conduct AI red-teaming, where testers deliberately try to make a system fail. The Office of Technology and Innovation disagreed with all seven.

How it ended

Mayor Zohran Mamdani announced the chatbot's shutdown in late January 2026, as a budget cut. He called it "functionally unusable." It was not shut down to fix its accuracy.

Reporting in 2024 corroborates a development cost of roughly $600,000. The figure of roughly $500,000 a year for maintenance, and the description of the chatbot as waste, come from the new administration.

As of early 2026, the chat.nyc.gov page reads: "The Chatbot beta test has ended."

Who bears the harm

The public asked the chatbot directly and could act on its answers. No caseworker or other professional stood between them.

The chatbot flagged no one and denied no one. It handed out instructions. Whoever followed a wrong one took on its legal risk. The harm could fall on people outside the conversation, such as workers and tenants, who had no channel to contest anything.

Why this case matters

Every model is wrong sometimes. The case file argues that the defining failure here was institutional, not technical. Reporters published the errors, an incident database listed them, and a formal audit documented them again. The city kept the chatbot online and planned to expand it, until a budget cut ended it. Exposure is not correction.

The case file contrasts this chatbot with two careful benefits chatbots in this Lab: Nava's assistive chatbot and the Imagine LA Benefit Navigator. In both, a caseworker checks the chatbot's answer before passing it on. Here the public asked directly. So there is no checking habit to protect, and no professional to train.

The case file points to four moves instead. Check each answer against an authoritative legal source before the public sees it. Label a generated answer as generated, and as outside what it may speak to. Agree in advance on a trigger that pauses a chatbot shown to be wrong. Make the outside audit binding rather than advisory.

What this network is drawn from

This network follows the pattern the case file describes. It does not reconstruct the actual chatbot. It shows the chatbot, the city's business pages, the public who asked, the disclaimer, and the outside reviewers. The workers and tenants who bear the risk are outside it.

A separate part of MyCity, a childcare benefits application flow, is not this chatbot. Its figures, including a roughly 46 percent application-ineligibility rate, belong to that flow. Being deemed ineligible is an application outcome, not proof of an algorithm's error.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Here a mistake is, for example, an answer telling a landlord he may refuse a housing voucher. A pathway counts as closed once it passes on only a few mistakes. The work along it goes on.

This case has a budget of 10 units. Each tool costs the same at every target level. Before any tool is used, four failure pathways are open. They are Answers to the public, Chatbot draws on the business pages, Public asks a question, and One chatbot answers everyone. Mistakes also build on one another across the network, so the targets are not met at any level that sets them.

Explore (No Targets) sets no targets. Under Service Targets Only, Escalate checks meets the targets on its own, for 2 units. When monitoring flags trouble, it raises how closely the chatbot's answers are checked, instead of waiting for the next review. On this network it has Business owners and the public check the chatbot's answers more closely. The case file argues a checking habit has no purchase here, because no professional stands between the chatbot and the person who asks. The sources describe no such checking by the public. Many other combinations within the budget also meet these targets. One is Mark AI-written records with Keep prompts neutral, for 4 units.

Under Service and Safety Targets and All Governance Targets, the targets can be met within the budget. Both levels ask you to close every failure pathway, among other targets. Each open pathway needs its own tool. Escalate checks closes Answers to the public. Mark AI-written records closes Chatbot draws on the business pages. Keep prompts neutral closes Public asks a question. Check with a second model closes One chatbot answers everyone.

Together those four cost 9 units. Every combination within the budget that meets the targets at these two levels includes all four. Counting each tool's settings, four such combinations meet them under Service and Safety Targets, and three under All Governance Targets.

Pause AI on alarms also closes Answers to the public, by halting the chatbot's answers until a person clears a review. But the chatbot then adds too little to the work. No combination within the budget that includes it meets the targets at any level that sets them.

More is not better here. Every tool at its strongest setting at once closes every failure pathway, but costs 34 units, more than three times the budget. It also leaves the chatbot adding too little to the work, so it misses the targets at every level that sets them.

Stylized model of a documented deploymentPublic benefits & eligibility

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Public-adviser-class generative chatbot with rejected external oversight network: 5 components and 9 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 4 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 8 assumptions
  • assumed

    This network follows the pattern the NYC MyCity Business chatbot case file describes: a public AI adviser answering in an official city voice. It does not reconstruct the actual chatbot. The case file sets it against two careful benefits chatbots in this Lab, Nava's assistive chatbot and the Imagine LA Benefit Navigator. There a caseworker checks each answer before passing it on. Here the public asks directly, so no professional checks first, and the harm falls on people outside the conversation.

  • baseline

    The network assumes the public acted directly on the chatbot's answers. The answers came in an official city voice, and no caseworker or other professional checked them first. The case file calls this the one design difference from the careful benefits chatbots, which use the same underlying technology.

  • baseline

    The network assumes no independent check tested an answer against the law or an authoritative rules source before the public saw it. That is how answers that would break the law reached business owners as apparent official guidance. They covered tips, landlords refusing tenants who pay with Section 8 vouchers or other income, cashless stores, and lockouts.

  • baseline

    One chatbot answered everyone, so the network assumes a wrong answer repeats for many people rather than scattering. When ten staffers at The Markup asked the housing-voucher question, all ten got the same wrong answer. The inconsistency the sources document is different: the same question got different answers over time, and after trivial edits to its wording.

  • assumed

    The network shows the beta disclaimer and the redirect patch as a named safeguard with no pathway of its own. The disclaimer told users to double-check through the links provided. The patch narrowed what the chatbot would answer. When asked, the chatbot itself still said it could be used for professional business advice. The double-check appears as the pathway named Public reads the linked city pages.

  • baseline

    The network assumes the audit's recommendations did not change the chatbot, because the city's technology office disagreed with all seven. One was AI red-teaming, where testers deliberately try to make the system fail. Reporters, an incident database, and the December 2025 Comptroller audit made the errors public. The chatbot ran for roughly two years, until a budget cut rather than an accuracy fix ended it. The case turns on this: a warning that arrives and is refused does not correct anything.

  • assumed

    How often the chatbot answered wrongly is the network's assumption, not a measured rate. The 2024 findings came from specific tested questions. The only figures come from the 2025 audit's sample and testing. They are 50 negative ratings from the 70 users who left feedback, and 23 of 48 tested government questions left unanswered. City officials claimed the chatbot had given thousands of people accurate answers, and "exponentially" fewer made-up answers, without figures.

  • assumed

    The people harmed are not in the network. Workers who could lose tips, and tenants who could be refused vouchers or locked out, were never in the conversation. They could bear the harm of a wrong answer. This Lab shows how mistakes pass between the parts of an organization. The case file describes where the risk falls outside it. The separate MyCity childcare application flow and its figures belong to a different system and are not shown here.

What this example does not show

Show all 3 limitations
  • This example does not show the people a wrong answer could harm. They include workers who could lose tips, and tenants who could be refused vouchers or locked out. They were never in the conversation. The Lab shows how mistakes pass inside an organization. The case file describes where the risk falls outside it.
  • The 2024 findings that the chatbot was confidently and repeatably wrong are qualitative. They rest on specific tested questions, not a sampled error rate. The only figures come from the December 2025 Comptroller audit's sample and testing. Of the 70 users who left feedback, 71.4 percent were negative (50 of 70). The city disputes that share, putting it at roughly 2.25 percent of all responses. The audit rebuts the city's way of counting. An internal weekly report reproduced in the audit showed 23 of 48 tested government questions unanswered. Neither figure is an error rate across every answer the chatbot gave.
  • This example shows the public advice chatbot only. The separate MyCity childcare application flow, and figures such as its roughly 46 percent application-ineligibility rate, belong to a different system. They are not shown or attributed to the chatbot here. The MyCity program's cost of over $100 million is for the whole system, not the chatbot alone.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • New York City launched the MyCity Business chatbot in 2023 on Microsoft Azure AI as a public-facing generative-AI adviser for business owners. A March 29, 2024 investigation by The Markup with THE CITY and Documented NY found it confidently and repeatably wrong on legal obligations, advising businesses in ways that would break the law, including that employers could take a cut of workers' tips, that landlords need not accept Section 8 vouchers or source-of-income tenants (illegal in New York City), that stores could go cashless against a 2020 city law, and that funeral-price disclosure could be concealed against the federal funeral rule; when ten staffers asked the housing-voucher question they received the same wrong answer, which had changed from an earlier correct one, showing the tool was non-deterministic. The 2024 findings are qualitative, based on specific tested questions rather than a sampled error rate. The city relabeled the tool a beta product with a disclaimer and applied a scope-narrowing patch rather than withdrawing it, kept it online for roughly two years, and shut it down in early 2026 as a budget cut rather than an accuracy fix.

    empirical
    • Investigative The Markup, NYC's AI chatbot tells businesses to break the law (2024); OECD AI incident https://themarkup.org/artificial-intelligence/2024/03/29/nycs-ai-chatbot-tells-businesses-to-break-the-law
    • Investigative The Markup and THE CITY, Malfunctioning NYC AI Chatbot Still Active Despite Widespread Evidence It's Encouraging Illegal Behavior (2024) https://themarkup.org/artificial-intelligence/2024/04/02/malfunctioning-nyc-ai-chatbot-still-active-despite-widespread-evidence-its-encouraging-illegal-behavior
    • Investigative Reuters (Jonathan Allen), New York City defends AI chatbot that advised entrepreneurs to break laws (2024) https://finance.yahoo.com/news/1-york-city-defends-ai-011454323.html
    • Investigative The Markup (Colin Lecher and Katie Honan), Mamdani to Kill the NYC AI Chatbot We Caught Telling Businesses to Break the Law (2026) https://themarkup.org/artificial-intelligence/2026/01/30/mamdani-to-kill-the-nyc-ai-chatbot-we-caught-telling-businesses-to-break-the-law
  • A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.

    empirical
    • Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
    • Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
  • A December 30, 2025 performance audit of the MyCity system, issued under New York City Comptroller Brad Lander, found the chatbot 'appears to be unable to provide accurate or consistent information' and reported that the wider MyCity system had cost over 100 million dollars across more than 120 agreements with about 50 vendors, lacked a system development plan, and had not delivered the promised single-form access to city benefits; the Office of Technology and Innovation disagreed with all seven of the audit's recommendations, including one to conduct AI red-teaming. Among the audit's figures, an internal weekly production report reproduced in the audit showed the chatbot did not answer 23 of 48 tested government questions, and of the more than 2,200 questions asked in July and August 2025 the 70 users who left thumbs-up-or-down feedback were 71.4 percent negative (50 of 70), a share the city disputes as roughly 2.25 percent of all responses, with the audit rebutting that denominator. The 100-million-dollar figure is the whole MyCity system, not the chatbot alone.

    empirical
    • Government evaluation Office of the New York City Comptroller (Brad Lander), Audit Report on the New York City Office of Technology and Innovation's MyCity System (2025) https://comptroller.nyc.gov/reports/audit-report-on-the-new-york-city-office-of-technology-and-innovations-mycity-system/

Where this connects

Institutional pressures in this domain

  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.

All of them in context on the Public benefits & eligibility domain page.

Levers available here and the patterns behind them

Documented case histories