Article · 19 August 2026
Why do AI chatbots get things wrong with such confidence?
Most chatbots built on large language models will give your customer a faulty answer with exactly the same confidence as a correct one. Looking at the answer, you cannot tell which one you just received. The tone, the flow and the wording are the same.
Tim Bisander
Founder & CPTO, Billuminate · 6 min read
The technology itself is not broken. The behaviour becomes understandable once you look at how a language model actually works. And if you understand how it works, you can do something about it.
How does a large language model work?
This is roughly how an LLM works. It starts by looking at the data you present it with, and answers based on that data are likely to be correct. But if your data is thin and does not cover the question, the model falls back on what it remembers from its training days. For common and stable knowledge, that usually works. For questions that are rare, market-specific or about circumstances that have changed, the model produces a confident answer based on probability. Sometimes it is right, sometimes it is not, and you cannot tell from the answer. Neither can the model.
A lot of relevant information changes over time: contact details, amounts, new laws. These are things the model has rarely seen, or that changed after its training ended. And since most models are trained overwhelmingly on English material, this matters even more in a smaller market like Sweden. Add one more thing: consumer-facing chatbots often prioritise answer speed over using the latest model, which in practice means a smaller model that remembers even fewer details.
The best medicine is more and better information. Give the model a maintained source to look the answer up in, and use the language model for what it is good at: the language.
What did we find in the Swedish credit industry?
During the summer we performed an extensive investigation into how public chatbots in the Swedish consumer credit industry answer real consumer questions, as part of an industry index we will publish shortly. A few patterns keep recurring, and none of them can be attributed to a specific vendor or tool.
01
Answers that look data-based but are made up.
Addresses, phone numbers and amounts are generated as language, not fetched as facts. The clearest error we saw: a chatbot that in several answers referred to a web page that does not exist and was never mentioned in the company’s own data, while other answers pointed to the real one. A fact fetched from a register cannot vary between answers. A fact written by the model can. The most unsettling instance was in the reply to a worried customer asking whether the company was real.
02
Answers delivered to the wrong audience.
We have seen consumer questions answered with great confidence out of a knowledge base written for corporate customers. We have also seen questions about interest rates on unsecured loans answered with mortgage information. The answer sounds relevant and the wording is correct, but it is based on the wrong information.
03
Answers based on outdated information.
We have seen faulty references to the Swedish amortization rules (changed on 1 April 2026) and a too low amount for the state deposit guarantee (raised to 1,150,000 kronor on 1 January 2026). The chat did everything technically correct and still gave the consumer the wrong answer.
Our investigation shows that most institutions are good at answering their own product-specific questions. There they seem to have solid quality control in place and review their content continuously. But we see big deviations between organisations in how well they answer questions outside their own product specifics. Which is understandable. It requires solid processes, systematic monitoring of changes, and resources.
Billuminate’s business is to provide companies with exactly this. We constantly search for changed information and identify legal updates, new credit bureau score models and similar changes early, before they reach the customer conversation. Naturally we find gaps in our own answers too. When we do, our processes flag them so that our experts can update the knowledge layer and verify the answer quality. That is how you constantly improve a system: measure, find, correct, and measure again.
1/3
On average, one third of the answers we tested were based on our verified knowledge layer. The rest is covered by the company’s own content, and that is how it should be.
4 of 10
But the deviations are huge. Some organisations can only answer four out of ten questions from their own data. The rest is filled in by the model’s own memory, with the consequences described above.
So what can you do? We like to follow three simple principles.
Three simple principles
The language model remembers approximately and does not know when it is filling a gap. You cannot train that away, but you can build around it.
01
Look it up, don't guess.
An AI that cannot find the answer in a designated source will pull it from its own training data, often outdated and shaped by American conditions. Use the model's memory for language, not for facts.
02
Watch for changing facts.
An answer that was right when written becomes silently wrong when the rule changes. Monitor changes to product rules, laws and the like, and date every update.
03
Measure the right thing.
Don't just ask how many cases the chat resolved. Ask whether the answers were true and current, and fix what was wrong or missing. Continuously. What isn't measured doesn't improve.
Whether you use us at Billuminate to do this for you, or decide to do it yourself, it needs to be done if you want to serve your consumer customers with correct answers.

About the author
Tim Bisander, founder and CPTO at Billuminate
The observations above come from our ongoing work on a quality index for public chatbots in the Swedish consumer credit industry. The index will be published on our website shortly.
LinkedIn →