To know whether an AI tool has answered correctly about your organisation's data, you need to see three things, the query the tool ran, the data that query returned and the definition used for each indicator. If the tool shows all three, any figure can be checked in a few minutes. If it only shows the final answer, you are trusting it blindly.
Natural language analytics tools have made it easy to ask "how much did we sell last quarter, by region?". The challenge has moved elsewhere. An answer written with confidence is not necessarily a correct one, and in a board meeting a wrong figure does more damage than a missing one.
Why can an AI get answers about company data wrong?
When someone asks a question in plain language, the tool turns it into a database query, usually in SQL, runs it and presents the result. An error can slip in at any of these stages.
- The question is ambiguous. "Sales" may mean revenue with or without VAT, with or without credit notes, counted by invoice date or by order date.
- The query reads the wrong table. Many ERP systems keep similar information in more than one place, for instance in draft documents and in final ones.
- The query is right but the filter is not. A badly defined period, a forgotten customer or a branch left out will change the total without raising any error.
- The written answer does not match the data. A language model can produce a convincing sentence quoting a figure that is not in the query result.
- The source data has problems. Duplicate records or empty fields produce a wrong figure even with a perfect query.
None of these errors is unique to artificial intelligence. A report built by hand in a spreadsheet runs the same risks. The difference is that AI answers quickly and with apparent certainty, so checking has to be designed into the tool itself rather than left to the scepticism of the reader.
How do you check an answer, step by step?
These five steps apply to any AI data analysis tool and take only a few minutes.
- Read the question as the tool understood it. A good answer states, in plain language, what the query returns. If that description is not what you wanted to know, the problem lies in the question, not in the data.
- Open the query. Look at which tables were read, which filters were applied and which period was used. You do not need to write SQL to notice a missing date filter, or that the table read holds quotes when you asked about invoices.
- Look at the returned data. The figure in the answer must appear in the results table. If the sentence mentions a value that is not in the data, be wary.
- Compare with a known figure. The total for a closed month, the revenue of a customer you know well or a number from an official report make a quick benchmark.
- Ask the same question another way. If "September revenue" and "total invoiced between 1 and 30 September" give different figures, there is a definition to clarify.
Why must the query and the returned data be visible?
A tool that only shows the final answer asks for an act of faith. A tool that shows the query and the returned data allows a check. That is the difference between a black box and a working instrument. In practice, transparency comes in three levels.
| Level | What you see | What it allows |
|---|---|---|
| Answer only | A figure or a sentence | Accept or reject, without knowing why |
| Answer and data | The table with the query result | Confirm the figure is in the data |
| Answer, data and query | The table and the query that produced it | Confirm where each figure comes from and rerun the query in another tool |
For financial or commercial decisions, only the third level is enough. It is also the only one that lets the people who manage your systems correct a query and reuse it. For a broader comparison of ways to query data, you can read our article on conversational analytics and dashboards.
How do you avoid right answers to the wrong question?
The most common error is not the AI inventing a figure. It is correctly answering a question slightly different from the one you meant to ask. This is solved more by organisation than by technology.
- Write down the definitions of your main indicators. What counts as a sale, an active customer or a late order. One page is enough.
- Always use the same reference period. Calendar month, fiscal month and the last 30 days give different figures, and all of them are right.
- Start with questions whose answers you already know. The first weeks are for building confidence in the tool with figures the team knows by heart.
- Keep the questions that work. A checked question can become a report or a dashboard, so it no longer needs checking every time.
Who can ask what, and what gets logged?
Trusting an answer also means knowing who asked for it and which data they could reach. Two questions help to assess any tool.
- Is access controlled at data level? Controlling who logs in is not enough. You need to define which tables each person can query, so that a question about salaries is not answered for someone who should not see them.
- Is each question logged? A log showing who asked, which source was used, which query ran and how many rows came back lets you trace where a figure presented in a meeting came from.
A practical example in a distribution company
Imagine a distribution company where the finance director asks "what was the margin by region in the third quarter?". The tool replies with a bar chart and a sentence saying the North region had the highest margin. Before taking the chart to the board meeting, the director opens the query and notices that credit notes were not deducted. They ask again, this time "margin by region in the third quarter, net of credit notes", and the North drops to second place.
In this example the tool did not make a mistake. It answered the question it was given. It was the visible query that showed the question needed to be more precise, before the wrong conclusion reached the meeting.
This is how EngiAnalytics, the Engibots conversational data analysis platform, works. Every answer comes with a plain language sentence describing what the query returns, and if that sentence quotes a figure that is not in the returned data, the reader is warned. The SQL query behind each answer can be opened and copied, next to the table with the returned data. Access profiles define, table by table, what each person can query, and a question outside that scope is refused at once. Every question is logged with who asked, the source, the query run, the number of rows returned, the duration and the outcome, and the returned data is not stored. Exploratory discovery scans the database for inconsistencies and shows what deserves fixing before it reaches a report.
Frequently asked questions
Can an AI invent figures about company data?
It can, especially in the text that accompanies the answer. That is why the tool should show the data returned by the query and warn the reader when the sentence quotes a figure that is not in that data.
Do you need to know SQL to check an answer?
No. It is enough to see which tables were read and which filters and period were applied. The people who manage your systems can copy the query and rerun it in another tool when confirmation is needed.
What is the most common error in AI answers about data?
Correctly answering a question slightly different from the intended one, such as sales including VAT when you wanted them excluding VAT. Written definitions of the main indicators prevent most of these cases.
How do you control what each person can ask?
With access profiles defined table by table. A question outside that scope should be refused without exposing restricted information.
What should be logged for each question?
Who asked, the source, the query run, the number of rows returned, the duration and the outcome. With that you can trace where any figure presented in a meeting came from.