AI Chatbots Get Financial Queries Wrong ‘Most Of Time’

Individuals could risk financial losses by heeding financial advice from AI chatbots, a fintech research firm has said, after its research found the most popular chatbots gave incorrect answers to financial queries in 57 percent of cases on average.

Accurate responses from the most popular offerings from AI chatbots such as OpenAI’s ChatGPT, Anthropic’s Claude, Microsoft’s Copilot, xAI’s Grok and Google’s Gemini gave accurate responses 43 percent of the time, on average, according to the study from Saturn.

Models made more mistakes when complex questions were involved, with the error rate rising to 88 percent on average, while some models provided incorrect information 99 percent of the time when asked harder questions.

Errors, hallucinations

Saturn tested 18 popular AI models against 121 financial questions, with each repeated 5 times to check consistency, for a total of more than 10,000 queries.

The responses included errors in calculations, missed important risk warnings, ignored upcoming tax changes, or invented rules that did not exist.

Free-to-use models provided less accurate advice, making mistakes 63 percent of the time on average, but paid models still made errors in an average of 49 percent of responses.

On the hardest questions, free models made mistakes in 93 percent of answers.

In one case, Claude Haiku 4.5, a free model, made a mistake on a pension tax question that could have resulted in a £17,500 charge from HMRC, while in another case a Claude model invented a student loan rule, saying wrongly that a person could stop repayments if they moved abroad.

“The low quality of financial advice from mainstream AI models risks leading to widespread consumer harm,” said Saturn chief executive Amal Jolly.

Research by the Financial Conduct Authority published last month found 26 percent of UK adults trusted general-purpose chatbots such as ChatGPT and Claude for financial advice.

The FCA has said it is considering whether financial advice from chatbots should be regulated.

Similar concerns have been expressed around the use of AI chatbots for health queries.

Recent research has found that AI models specialising in healthcare also experience high failure rates, with a recent study in finding that none of the models it examined, including ChatGPT Health, were ready for deployment, in findings that OpenAI said were inaccurate.

Original source AI Chatbots Get Financial Queries Wrong ‘Most Of Time’

Back to home