AI hallucination, data privacy and human approval in projects
Three risks come up most often in AI projects: the model confidently producing false information (hallucination), sensitive data ending up in the wrong place, and the system taking an irreversible step without a person’s approval. None of the three can be removed entirely — but handled from the first day of design, each becomes small, visible and manageable. In this article we show, with concrete examples, how we deal with each of them.
Why “we’ll sort it out later” does not work
Most AI projects begin with a demo. The demo is impressive: the model answers questions, summarises documents, writes drafts. The trouble starts when real users, real data and real consequences arrive.
At that point, adding security afterwards is expensive. Which data goes to the model, which step needs approval, which decision gets recorded — these are load-bearing parts of the architecture. Bolt them on later and you end up rewriting half the system. So we answer these three questions in the project’s first written note.
Hallucination: why does a model make things up?
Hallucination is when a language model presents something that is not true as if it were fact. A regulation clause that does not exist, a wrong date, an invented source.
The cause lies in how the model works. A language model is not a database of facts; it produces the most likely continuation of a piece of text. Even on a subject it knows nothing about, it can produce an answer that sounds fluent and reasonable. Fluency is not accuracy.
Ways to reduce hallucination
You cannot bring it to zero. But you can reduce it markedly:
- Ground the answer in sources. Have the model answer from the company’s own documents rather than its general memory. This is the basis of RAG (retrieval-augmented generation).
- Show the source. Show which document each answer rests on. The user can check it, and the system can check itself.
- Teach it to say “I don’t know”. Instruct the model to say plainly when the answer is not in the sources — and test that it does.
- Narrow the task. Rather than “answer everything”, give a narrow brief such as “answer only questions about the returns policy”.
- Structured output. Ask for output with defined fields instead of free text, and validate those fields in code. If anything other than a date lands in a date field, the system catches it.
- Measure. Build a test set from real questions and run it after every change.
An example
Someone asks an insurance broker’s assistant: “Does my car policy cover earthquake damage?” A poorly built system might answer from general knowledge: “Yes, comprehensive car policies usually include earthquake cover.” That may well be wrong for this customer’s policy.
A well-built system looks at the customer’s actual policy document. If it finds earthquake in the list of covers, it answers and shows the relevant clause. If it does not, it says: “I could not find a clause on earthquake cover in your policy; for a definite answer, I would suggest speaking to your broker.” The second answer is less impressive — but it is correct.
Data privacy: where does each piece of data go?
Every piece of text sent to an AI system goes somewhere. Where that is, how long it stays there and who can reach it are questions to answer at the very start of a project.
What we pay attention to:
- A data inventory. Listing the data the system will touch: customer names, contact details, financial data, health information, trade secrets.
- Minimum data. Sending the model only what the task needs. Classifying a support request does not require the customer’s national ID number.
- Masking. Where personal data is not needed, masking it or replacing it with a pseudonym before it reaches the model.
- Where it is processed, and the contract. Knowing in writing where the model runs, how the provider stores what is sent, and whether it is used for training.
- Permission boundaries. The assistant must not reach data that the person using it could not already reach.
- The logs themselves. System logs contain data too. How long they are kept, and who can see them, is decided separately.
Projects that handle personal data fall under data-protection law — GDPR in Europe, or KVKK, Türkiye’s data-protection law and broadly similar to GDPR — and fields such as health and finance add sector-specific rules on top. We leave the legal assessment to legal specialists and build the technical side to match.
Human approval: who owns which step?
Human approval means an AI system waits for a person to approve certain steps before it takes them. Put everything behind approval and the system becomes useless. Put nothing behind it and it becomes risky. The right balance comes from classifying the steps.
| Type of step | Example | Our approach |
|---|---|---|
| Read | Looking up a record, searching documents | No approval, but logged |
| Draft | An email draft, a suggested note | Drafted freely; sending needs approval |
| Reversible action | Adding a tag, opening a ticket | Depending on context, no approval but limited |
| Irreversible action | Payment, deletion, sending outside | Always needs approval |
We fill in this table with the client on every project. Which step goes into which row is not a technical decision — it is a business decision.
Kill switches and limits
Alongside human approval, we build two more mechanisms:
- A kill switch. The ability to take the system out of service instantly, from a single point. In CAI, our own autonomous trading system, manual override always wins over automatic decisions, in one click. Building that in from the start was the precondition for having the confidence to run the system with real capital.
- Budget and step limits. How much an agent may spend, how many steps it may take, and when it must stop and ask a person. A system left without limits becomes unpredictable — in cost and in errors alike.
Logging: the answer to “why?”
When something goes wrong, the first question is always “why?”. That question can only be answered if there is a record. We log the inputs to every decision, the sources used, the model’s output and who approved it.
Logs are not only for debugging; they are for trust. Being able to show a customer, an auditor or your own team how the system reached a decision stops AI from being a black box. We describe how this works in a system of many cooperating agents in our article on multi-agent AI systems.
The questions we ask in a first conversation
When we start an AI project, we put these three risks on the table with a short list of questions:
- If the system gives a wrong answer, what is the worst that could happen? Who is affected?
- Which data will the system touch? Is any of it personally or commercially sensitive?
- Which model will that data go to — and to servers in which country?
- Which steps cannot be undone? Who will approve them?
- Who can stop the system, from where, and how quickly?
- When an error happens, how will we see its cause?
- How will we judge success?
The answers form the skeleton of the project’s written architecture note. Every question left without a clear answer is an early warning of a future risk.
Frequently asked questions
Can hallucination be prevented completely?
No. With today’s language models, zero errors is not possible. But grounding answers in sources, narrowing the task, teaching the model to say “I don’t know” and measuring regularly bring the error rate down markedly. We manage the remaining risk with human approval and logging.
Will our data be used to train the model provider’s models?
That depends on the provider and on the terms of the service used. At the start of the project we set out in writing which provider will be used and on what terms, and we look at those terms especially closely for work involving sensitive data.
Isn’t human approval on every step the safest option?
On paper, yes; in practice, no. In a system where everything needs approval, approval soon becomes a click given without reading. Keeping approval for the steps that genuinely matter keeps it meaningful.
Can these controls be added to an existing AI project?
They can, but it usually means changes to the architecture. We start by reviewing the existing system and writing a short architecture note that ranks the risks; the most critical gaps are closed first.
If you would like to go through the risks of your own AI project together, have a look at our AI systems architecture page or write to us. We reply within two business days.