Skip to content

PAN Lab example

Singapore's chatbot fleet refresh

Eighty engines into one: a whole-of-government chatbot fleet refresh

Singapore set out to move its agencies' chatbots off scripted engines, which match keywords to prewritten answers, onto a few shared AI engines.

See more

VICA, the Virtual Intelligent Chat Assistant, is a chatbot platform run by GovTech, Singapore's Government Technology Agency. Each agency's chatbot keeps its own branding and its own bank of prepared questions and answers. VICA's shared AI engines answer the questions for all of them, so one engine mistake can show up at many agencies at once.

How it began

Singapore's government chatbots began in 2014 with Ask Jamie. A survey had found that roughly half of visitors' questions to government agencies were general enquiries.

Ask Jamie was scripted. It matched keywords against each agency's own curated question-and-answer bank. It passed complex questions to human channels with the conversation history kept. Each agency website ran its own copy over its own content.

It also handled some transactions through SingPass, such as Central Provident Fund statement opt-ins and tax-filing status checks. A voice version was piloted with the Ministry of Social and Family Development and the Ministry of Education.

What the vendor claims

Ask Jamie's vendor, Sabio, formerly flexAnswer, makes claims that are not official statistics. It says that about five years after launch, Ask Jamie ran on 80 government websites and 9 intranet sites. It says the chatbots had answered over 15 million questions since 2014. It says they cut enquiries that would have gone to call centres by up to 50 percent. It cites a whole-of-government "no wrong door" policy, which the sources do not describe further. A later trade-press count found over 70 sites.

One agency's failure, one agency's fix

In October 2021 the Ministry of Health temporarily switched off its own Ask Jamie chatbot. Residents had shared misaligned COVID-19 answers online. A question about a daughter testing positive drew safe-sex advice. An antigen rapid-test question drew information about the polio vaccine.

The ministry said it had disabled the function "to allow us to conduct a thorough system check and work on improvements." It sent users to covid.gov.sg. Every other agency's Ask Jamie ran on untouched.

The move to shared engines

In 2023 GovTech began refreshing government chatbots under the VICA project. This case calls all the government's chatbots, taken together, the fleet. GovTech set out to move them from scripted engines to centrally provided large-language-model engines.

GovTech aimed to convert all 88 government chatbots and retire the scripted Ask Jamie engine by the end of 2023. The verified snapshot is 21 of 88 converted, as of September 2023 reporting. No independent source confirms the migration finished on schedule.

The move spans government websites, WhatsApp, Telegram channels, and internal portals. The engines reportedly run on commercial cloud large-language-model services. The reporting names them, and this case leaves them out.

The platform in 2026

By GovTech's VICA product page, last updated 29 April 2026, the platform hosts over 100 chatbots for more than 60 agencies. About one fifth serve government staff. It averages over 800,000 questions a month. These are the government's own figures.

The page describes Hybrid AI, which combines predictable language processing with generative AI. Its VICA 2 lets agencies "ground, steer, and override" generative answers. They use public information sources and preferred question-and-answer sets.

The safeguards GovTech described

Keeping a human in the loop is one key way GovTech keeps the chatbots trustworthy, GovInsider reported in 2023. The one review it describes is of generated question-and-answer pairs, before agencies add them to their banks.

Those pairs would come from a planned automated generator. It would extract them from government documents, such as PDFs and websites. The sources do not confirm it shipped.

GovTech also planned a scoring system to help agencies find inaccurate or unhelpful answers. The chatbots would also cluster questions they cannot answer, so agencies know what to add. No source located confirms either now runs.

Nothing here decides eligibility

The SupportGoWhere benefits hub grew out of GovTech's COVID-era GoWhere suite. That suite had 16 initiatives with more than 34 million website visits as of January 2022. On 14 September 2023, Minister Josephine Teo announced a pilot of large language models in the hub's Support Recommender, which suggests support schemes. Citizens could describe their needs in their own words instead of filling in forms.

The Ministry of Finance's Support For You Calculator turns details people enter into estimates of Budget benefits. Its outputs are explicitly estimates, not entitlement decisions.

Chat.Gov.SG (Beta) "uses AI to summarise information from official government websites," its explainer says. It states it does not assess eligibility, make decisions, submit applications, or complete transactions. It warns users not to share personal or sensitive information, and does not promise fully accurate translations. The explainer's file details credit the Public Service Division and a creation date of 30 April 2026.

So a chatbot mistake here writes to no entitlement record. The harm is misdirection: a wrong scheme, a wrong agency, a wrong deadline, or a benefit never claimed. It is not a wrongful denial.

The trade the move made

The scripted chatbots failed one agency at a time, and could be fixed the same way. With shared engines, one fault can be wrong the same way at every agency at once. But fixes, guardrails, and review tools also go out to every agency in one move.

The case file calls this a trade: failures fixed locally for failures shared and fixed centrally. The case asks what sharing the engine gained, and what watching it would take.

What the sources do not show

No source located gives an accuracy or error rate for either the scripted or the shared-engine chatbots. None gives override or escalation counts, a before-and-after evaluation of the migration, or an independent or external audit.

The documented controls are a person reviewing generated question-and-answer pairs and the planned scoring system. The incident record is the single 2021 suspension, before the migration. Singapore publishes far less adversarial reporting than the heavily litigated systems in other cases here. So the absence of further incidents is not evidence the chatbots made no mistakes.

Where the facts come from

The facts come from a 2023 GovInsider report on the refresh and a 2019 GovTech feature on Ask Jamie. They also come from Sabio's vendor case study and a 2021 Marketing-Interactive report on the Ministry of Health suspension. Government sources are GovTech's VICA product page, the Chat.Gov.SG explainer, a 2022 GoWhere factsheet, and Minister Teo's 2023 speech. A 2024 Mustsharenews report describes the calculator.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other. A closed pathway is one that mistakes stop passing along. The work along it may go on.

This case has a budget of 11 units. Contained means the network's mistakes are corrected rather than building on each other.

Each tool has a standard and a stronger setting. The Lab opens with side-effects on, where governance carries its own costs. It also opens with lingering effects on, where damage outlasts its cause.

Under Explore (No Targets), which sets no targets, one tool alone keeps the mistakes contained: Mark AI-written records, at 2 of the 11 units. Gate record entries alone does it for 3 units.

Store less data alone does it at its stronger setting, for 5 units. With lingering effects off, its standard setting does it for 3.

Under Service Targets Only, each of those also meets the targets. That level also asks for the automated system to be helping the work. Many other combinations meet them too. Escalate checks with Peer sharing rules does it for 4 units.

Under Service and Safety Targets and All Governance Targets, you must also close every failure pathway, among other targets. Counting stronger settings, two combinations meet the targets there, and each costs all 11 units.

Both combine Escalate checks, Peer sharing rules, Mark AI-written records, and Keep prompts neutral. The fifth tool is either Store less data or Gate record entries. Both leave the automated system clearly helping the work.

Each of the first four closes pathways no other tool here closes. Escalate checks closes the engine's answers to agency teams. Peer sharing rules closes GovTech's fixes to every agency. Mark AI-written records closes the reads of the banks and pages.

Keep prompts neutral closes GovTech's settings shaping the engine. Store less data and Gate record entries each close the writes into the banks and pages.

The three tools the case's argument centres on are in neither combination. Check with a second model leaves mistakes building on each other more, with side-effects on. Review on schedule and Gate vendor updates change no pathway on this network.

Stylized model of a documented deploymentBenefits navigation & public-facing chat

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Singapore-fleet-refresh-class shared-engine navigation platform network: 6 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 2 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 5 assumptions
  • assumed

    This example follows the move to shared chatbot engines described in the Singapore case file. It does not rebuild the actual platform. Its subject is how one decision changed whether the same mistake appears at many agencies at once. Scripted chatbots on more than 70 agency websites, 80 by the vendor's claim, were set to give way to a few shared central engines. No accuracy figure has been published for either period. So every assumption here about how often mistakes happen is a modeling choice, not a measured rate.

  • baseline

    Before the move, each agency ran its own chatbot over its own question-and-answer bank. A failure stayed with one agency, which could switch its chatbot off alone. The Ministry of Health did so in 2021, while every other agency's chatbot ran on. After the move, one fault can hit many agencies at once. It might be a new flaw in how the engine draws on its sources, a hidden-instruction attack, or a supplier's update. Here such a fault passes along the engine's answers to agency teams and the engine's use of the banks and pages. GovTech's fixes, guardrails, and review tools also apply to every agency at once. The move swapped failures fixed locally for failures shared and fixed centrally. This example asks which of the two you must watch for.

  • baseline

    The network places each safeguard where the sources put it. The sources document two controls: a person reviewing generated question-and-answer pairs, and a scoring system GovTech planned in 2023. The network draws the scoring system, with its clustering of questions, as a check. It is an internal check, GovTech's tools for its own chatbots. The network also holds two checks the sources do not document. One is a different engine checking the shared one, so a shared failure shows up as disagreement. The other is an independent evaluation across all agencies, with published figures. No published accuracy figures, no before-and-after evaluation of the migration, and no external or independent oversight were located. The new platform's performance figures are all the government's own. The Ask Jamie figures, 80 websites, over 15 million answers, and up to a 50 percent cut in call-centre enquiries, are vendor claims. The vendor's 80 websites differ from the trade-press count of over 70.

  • baseline

    The chatbots decide no one's eligibility. The shared engine points people to schemes, explains them, and hands them on to official portals. The calculator gives estimates that are explicitly not entitlement decisions. The Chat.Gov.SG (Beta) explainer states it does not assess eligibility, make decisions, submit applications, or complete transactions. It also warns users not to share personal or sensitive information. So a shared-engine mistake sends the wrong scheme, agency, or deadline to people across government at once. Or it leads to a benefit never claimed. This example tracks that as a pattern inside government, never as an outcome for any person. That limit on scope is why privacy exposure here is modest. It is also why the leverage lies in whether the same mistake appears at many agencies at once, not in a decision step.

  • assumed

    The citizens who use the chatbots are not part of how this example behaves. It tracks patterns inside government only. The operators describe handling Singlish, the local informal English, and restating answers in several languages as inclusion features. Those are not measured differences in harm, and this example computes no outcome for any group. No incident has been documented since the 2021 suspension. Singapore publishes far less adversarial reporting than the heavily litigated systems in other cases here. So that silence is not evidence the chatbots made no mistakes. A misdirection, a rerouting, or a gap in evaluation here is a pattern inside government, never a person. This example does not promise that any real deployment is safe.

What this example does not show

Show all 4 limitations
  • The citizens who use the chatbots, and the benefits they do or do not go on to claim, are not shown in this example. It shows how mistakes move between the engines, agency teams, GovTech, and the records. The chatbots decide no one's eligibility. So the harm this example tracks is misdirection, such as a wrong scheme, agency, or deadline, or a benefit never claimed. It is never a wrongful denial, and no outcome for any person is computed.
  • No public source located gives accuracy figures, a before-and-after evaluation of the migration, or override and escalation counts, for either the scripted or the shared-engine chatbots. So every assumption here about mistakes is a modeling choice, not a measured rate. The new platform's figures are the government's own: over 100 chatbots, more than 60 agencies, and over 800,000 questions a month on average. The Ask Jamie figures are vendor claims: 80 websites, over 15 million answers, and up to a 50 percent cut in call-centre enquiries. They differ from a trade-press count of over 70 sites.
  • No independent source confirms the migration finished. The sources show an end-2023 target, and 21 of 88 chatbots converted as of September 2023. A 2026 product page implies, but does not state, that every chatbot moved. No incident has been documented since the Ministry of Health's 2021 suspension. Singapore publishes little adversarial reporting, so that silence is not evidence the chatbots made no mistakes.
  • This example does not promise that any real deployment is safe. Two checks in the network are ones the sources do not document: a different engine checking the shared one, and an independent evaluation across all the government's chatbots. They stand for the oversight a governed move to shared engines would build. No source located shows either was built.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • In 2023 Singapore's GovTech began a whole-of-government retire-and-replace of its scripted Ask Jamie chatbots, embedded since 2014 on 70-plus (a vendor case study claims 80) agency websites as independent per-agency answer engines, migrating government chatbots onto centrally provided large-language-model engines; the stated aim was to convert all 88 chatbots and retire the scripted engine by end 2023, the verified snapshot is 21 of 88 converted as of September 2023 (migration completion not independently documented), and by the VICA product page updated 29 April 2026 the successor platform hosts over 100 chatbots for 60-plus agencies at an average of over 800,000 monthly queries, figures that are all government self-reported.

    empirical
    • Trade press Hirdaramani, Is it time to say goodbye to Ask Jamie? Inside GovTech's refresh of government chatbots (GovInsider, 2023) https://govinsider.asia/intl-en/article/is-it-time-to-say-goodbye-to-ask-jamie-inside-govtechs-refresh-of-government-chatbots
    • Government GovTech Singapore, Virtual Intelligent Chat Assistant (VICA) product page (2026) https://www.tech.gov.sg/products-and-services/for-government-agencies/informational-services/vica/
    • Government GovTech Singapore, Get to know the GovTech team behind Ask Jamie, the government chatbot (2019) https://www.tech.gov.sg/technews/govtech-team-behind-ask-jamie-government-chatbot/
  • The Singapore government benefits-navigation surface is documented as scope-limited to information and estimates rather than adjudication: the Ministry of Finance Support For You Calculator turns self-declared inputs into benefit estimates that are explicitly estimates and not entitlement decisions, and the Chat.Gov.SG (Beta) explainer hosted on the SupportGoWhere domain states the assistant summarises information from official government websites and does not assess eligibility, make decisions, submit applications, or complete transactions, and warns users not to share personal or sensitive information.

    empirical
    • Government Public Service Division (Singapore Government), About Chat.Gov.SG (Beta) explainer (2026) https://supportgowhere.life.gov.sg/learn-more-about-sgw-chatbot.pdf
    • Trade press Mustsharenews, Budget 2024 online calculator helps work out how much you stand to benefit (2024) https://mustsharenews.com/budget-2024-calculator/

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Benefits navigation & public-facing chat domain page.

Levers available here and the patterns behind them

Documented case histories