Rayify

Rayify builds decision intelligence systems for financial services and trains the teams who run them.

The Trust Deficit: Why AI-Enabled Financial Advice Needs Better Regulatory Oversight

Nearly half of consumers already ask AI what to do with their money, yet the supervisory toolkit still assumes a human is giving the advice. Here is where oversight has to go next.

Published

# The Trust Deficit: Why AI-Enabled Financial Advice Needs Better Regulatory Oversight

Nearly half of consumers already ask an AI what to do with their money. The oversight built to protect them still assumes a human is on the other end of the conversation. At Rayify we spend our days stress-testing how automated systems reach a recommendation, and the gap between how fast people trust AI advice and how well anyone can check it is the widest we have seen in a regulated market. This post is about that gap - why AI financial advice fails in ways that are hard to see, why the current supervisory toolkit cannot catch those failures, and where we think oversight has to go next.

## The adoption is already here

The debate about whether people should use AI for financial decisions is over; they already do. EY's 2026 survey of more than 18,000 consumers across 23 countries found that nearly half of global consumers now use AI to guide savings and investment decisions, a figure that rises to 68% among Gen Z. The most-used channel is not a regulated robo-adviser but a general-purpose chatbot: OpenAI reports that more than 200 million people come to ChatGPT every month for budgeting and questions about their investments.

So the question is not whether AI shapes financial decisions. It is whether we can trust *how* that advice is generated. And here the honest answer is that confidence is not competence. A model can produce a fluent, personalised pension recommendation that reads exactly like expert guidance and is wrong in ways neither the consumer nor the firm can detect.

## The hard part is not hallucination in isolation

It is tempting to reduce the problem to "AI hallucinates" or "AI is biased." Both are true, but the harder issue is that these systems fail quietly. Four failure modes matter more than raw error rates because each one is difficult to observe from the outside:

- **They assume instead of asking.** A suitability conversation is supposed to gather a client's circumstances before recommending anything. A model will often fill the gaps itself, inventing an investment horizon or risk appetite the client never stated, and then produce a recommendation that looks suitable against facts that were never true.
- **They skip compliance checks silently.** Nothing in a fluent answer tells you whether a required step was performed. The output looks the same whether the model checked a rule or ignored it.
- **They produce reasoning theatre.** The explanation a model gives for its answer often does not reflect the computation that produced it. Research shows that language models do not always say what they think, generating plausible chain-of-thought rationales that systematically omit the real influences on their answer. For a regulator that relies on an audit trail, an explanation that is not faithful to the decision is worse than no explanation at all.
- **They answer the same question differently on phrasing alone.** Language models are sensitive to spurious features in prompt design: a trivial change in formatting or wording can swing an answer. Two clients with identical circumstances who phrase the same question slightly differently can receive materially different suitability assessments.

None of these is visible in a single transcript that looks polished. That is precisely why they are dangerous at scale.

## The problem space regulators now face

LLM agents are no longer confined to FAQ chatbots. They handle product discovery, run suitability-style conversations, and field questions on pensions, insurance, and debt at a volume no human network could match. Meanwhile the unregulated foundation model is frequently the *first* place a consumer turns, before any authorised firm is in the loop.

Regulators have noticed. The FCA has published the Mills Review into the long-term impact of AI on retail financial services, the first review of its kind commissioned by a financial regulator, which frames AI as moving from a back-office efficiency tool into a more autonomous layer of regulated service delivery. The direction of travel is clear: the highest-volume advice channel may sit at or beyond the edge of the regulatory perimeter.

## The risks, concretely

**Unreliable advice at scale.** A human advice network fails one client at a time. A single model failing is a single point of failure that no human network ever had: one flawed system gives the same bad advice to millions of people simultaneously, and it does so without any of the local judgement that catches an outlier in a human process.

**Hallucinated product facts.** Professional advisers reviewing AI-generated portfolios have found them heavily US-centric and poorly tailored to UK rules such as ISAs and gilts - plausible, confident, and wrong for the consumer in front of them. Broader testing shows that AI programs give inconsistent, inaccurate, and sometimes biased recommendations on emergency savings, asset allocation, and retirement withdrawals.

**Advice-like outputs with no protections.** When a general-purpose tool produces something that walks and talks like regulated advice, none of the regulated advice protections attach. There is no suitability duty, no suitability report, no Consumer Duty obligation to deliver good outcomes, and no route to the Financial Ombudsman Service if the advice causes harm.

**Accountability gaps and AI-washing.** Modern agentic advice runs on a supply chain: a foundation model from one vendor, a retrieval layer from another, a fine-tune and an orchestration layer from the firm. The Senior Managers and Certification Regime pins accountability on a named individual, but that model strains when the decision-maker is a black box assembled from components no single person fully controls.

## Where oversight goes next - our thesis

The answer is not a smarter model. It is better mechanisms to test, audit, and explain how these systems behave. That is the problem Rayify works on, and it is why we think the next generation of supervisory tooling looks less like a rulebook and more like a test harness. We have been prototyping in each of these directions: an advice-market radar that maps where AI advice is actually consumed; a financial-advice red-teaming engine that probes advice systems with adversarial synthetic personas; a cross-market benchmark for agentic advice; a consumer explainability translator; and a supervisory explainability toolkit.

This is the same discipline we bring to Rayify's own product: we stress-test advice with synthetic personas before it reaches a person, and we translate a model's decision into an explanation that a regulator or a consumer can audit rather than merely admire. Trust in AI-enabled financial advice will not come from hoping the models are right. It will come from building the mechanisms that let us check.

All Insights