Lesson 21 of 64 Module 4: When it is wrong
Why AI makes things up
After this lesson you can
- Explain why a confident answer can be completely invented.
- Predict which of your questions are high risk.
- Ask in a way that lets it admit it does not know.
- Stop taking confidence as evidence.
Read first: It is guessing the next word
You ask an AI app how much money one scheme gives in your block. It answers in one second, with an exact figure and a last date. Next week the clerk at the office tells you no such amount exists.
The app was not playing a trick on you. It was doing the thing it was built to do.
An exam with no negative marking
There is a name for this. When a model invents something and states it plainly, that is called a hallucination. It is an odd word, but it is the one everybody uses.
It is not lying. A lie needs somebody who knows the truth and hides it. The model does not know its answer is wrong. There is no store of facts inside it to check against.
Now think about the exams you know. In JEE, NEET and GATE, a wrong answer cuts your marks. So a careful student leaves a hard question blank.
Picture the opposite exam instead. No negative marking at all. A blank scores zero. A guess might get lucky. Every sensible student writes something in every box.
These models were graded that way while they were built. Their tests scored each answer right or wrong. Saying “I do not know” scored the same as being wrong, which is nothing. So guessing became the habit, and the habit stayed.
You can see the result in how rarely they refuse. In a large study of free assistants in 2025, more than three thousand questions were put to them. They declined to answer only about one time in two hundred.
A made up answer looks exactly like a real one
This is the whole problem, in one line. There is no tone of voice to listen for.
An invented answer arrives in the same calm, tidy English as a correct one. Same speed. Same confidence. Often the same neat formatting.
It can even sound more certain than the truth deserves. A person who half remembers a date says “sometime in autumn”. A model that is guessing says “30 September”. Being exact is not proof of knowing. Sometimes being exact is a sign of guessing.
Researchers tested this in the simplest way possible. They asked one model for the same person’s birthday three separate times. They got three different wrong dates.
None of those dates was looked up anywhere. As it is guessing the next word explains, the machine writes an answer one small piece at a time. A true date and an invented date fit the sentence equally well.
Which of your questions are the risky ones
You cannot check everything you ask. You do not need to. You need to know which questions are dangerous before you send them.
The risky ones are rare, local, specific and one-off. Your village pincode. One scheme’s exact amount this year. The dates in one officer’s career. The fee at a small private college in your district.
Think about why that is. The model learned patterns from an enormous amount of writing. A fact that appears once in all of that leaves no pattern behind. Researchers put a floor on this. Imagine two facts in ten of one kind appear only once in all that writing. Then the model will get at least two in ten of them wrong.
The safe questions are the opposite. How photosynthesis works. What a percentage is. How to write a leave letter to a headmaster. Those have been written about thousands of times, in much the same way, for years.
| Your question | Risk | Why |
|---|---|---|
| What is a percentage | Low | Explained the same way many thousands of times |
| How photosynthesis works | Low | Sits in every school book |
| Fee at one small college | High | Rare, local, and it changes every year |
| Exact amount of one scheme | High | Specific, local and time bound |
| Dates in one officer’s career | High | Appears once in all that writing, or never |
Indian questions carry extra risk on top of this. Why it is weaker on Indian questions covers that on its own.
Give it the paper instead of asking from memory
You met this idea earlier in the course. It is the single most useful thing in this whole lesson. Do not ask the model to remember a document. Hand it the document.
Paste the notice into the chat. Or take a clear photo of the page. Then ask your question about that text only.
The gap is measured, and it is wide. Give a model a document and tell it to stick to that document. The best ones then went against it in roughly two to five answers in a hundred. That was the picture in May 2026. Working from memory alone, the same kind of model gets things wrong far more often.
Here is the shape of the request.
I am pasting a notice below. Answer only from this text.
If the answer is not in this text, write NOT IN THE TEXT.
Do not use anything you know from outside it.
[paste the notice here]
Question: what is the last date to apply, and who can apply?
Be careful with that low number, though. It is for the easy job of sticking to a page you handed over. It is not a score for how often the app is right in general.
Giving it your own book, notes or photo shows how to do this on a phone.
One line that lets it say it is not sure
The guessing habit came from grading. You can push back against that grading yourself, in one line, on every factual question.
Answer only if you are sure.
If you are not sure, write NOT SURE instead of guessing.
For each fact, tell me whether you are sure or not sure.
A useful answer then looks like this.
Runs under the state government: sure.
Last date this year: NOT SURE. Check the department website.
Exact amount for your block: NOT SURE.
That is a good outcome, not a failure. It tells you which two things to go and confirm, and which one you can carry.
Be honest about what the line does. It changes what the model is trying to score well on. It does not hand it knowledge it never had. The researchers who studied the grading problem suggested exactly this kind of instruction. It costs you one line, so use it every time.
Search helps, but it does not end the job
Many apps can look things up on the web while they answer. Switch that on for factual questions. It genuinely lowers the invention.
It does not finish the job. A hard set of questions was tested in early 2026. Even the strongest setup, with web search on, still made up claims in about three answers out of ten. The links looked correct. The pages did not always say what the answer claimed.
So search moves the work rather than removing it. Your question stops being “is this true?” and becomes “does that page really say this?”. Only you can answer the second one. Checking is a habit, not a setting you switch on once.
One more thing decides how much invention you face, and that is the tool itself. Smaller and older models make things up far more. In one company’s own tests in 2025, its small cheap model was wrong on about eight out of ten short factual questions. Its larger model was wrong on about half of them.
So a very old free app is not a bargain. It is a worse tool. How it searches, sees and uses tools shows what really happens when an app looks something up.
What it costs when nobody checks
This is not a small worry for careless people. In India, trained professionals have been caught filing documents that quoted judgments which do not exist. The filings were thrown out. The people who filed them paid for it, in money and in standing.
If a lawyer can be fooled by an invented reference, so can you. The rule that profession drew from it works for everyone. Never quote something you have not opened and read yourself.
Researchers who track such court cases around the world counted well over a thousand by the middle of 2026. The count keeps climbing as more people use these tools without checking.
Fake references are convincing because they are built to look ordinary. Checking sources, links and quotes deals with that in full.
Where this goes next
You now know why it invents, and which of your questions pull the invention out. That is half the skill.
The other half is fast checking. The next lesson, the two-chat test and three other checks, gives you four checks that take under two minutes.
Do this now
Ask about your own block, twice
- Pick one small thing near you. Your block office, a local college, or one scheme you have heard of.
- Ask the app for an exact detail. An amount, a last date, a fee or a phone number.
- Write the answer on paper, word for word, including every number.
- Open a brand new chat. Ask the same question in the same words.
- Compare the two answers. If any detail changed, neither one is knowledge.
- Now ask a third time and add the line: if you are not sure, write NOT SURE. Note what it admits.
Remember this much
- Inventing an answer is called a hallucination. It is not a lie, because it does not know.
- It was graded like a student in an exam with no negative marking, so it guesses.
- A made up answer looks and sounds exactly like a correct one.
- Rare, local, one-off facts are the dangerous ones. Common school topics are not.
- Give it the paper, and add the line: if you are not sure, write NOT SURE.
- Web search lowers the invention. It does not end your checking.
Questions people ask
What is AI hallucination in simple words?
It is when an AI app invents something and says it as if it were true. The word sounds dramatic, but the behaviour is ordinary. It happens most on rare, local and very specific facts.
Why does ChatGPT give wrong answers so confidently?
Because confidence is how it writes everything. It has no separate feeling of knowing or not knowing. The tests used to grade these models gave nothing for saying I do not know, so guessing became normal.
Does asking are you sure fix a wrong answer?
No, and it can make things worse. That question changes the words it writes, not what it knows. It may also drop a correct answer just because you pushed. Check a real source instead.
Can AI invent a website link or a book name?
Yes, and invented names are built to look real. Researchers found that a small share of software names these tools produce do not exist at all. Never type a link from an AI into your browser without checking it opens the real site.
Will newer AI versions stop making things up?
Bigger and newer models invent less than small old ones, so the tool you pick matters. But the habit comes from how these models are graded, not from a small bug. Plan to keep checking.
Prices, free limits and app screens change often. The facts in this lesson were checked on 9 August 2026. If what you see on your phone looks different, trust your phone and read the idea, not the exact button name.