Lesson 22 of 64 Module 4: When it is wrong

The two-chat test and three other checks

9 min read Free, no sign-up 9 August 2026

After this lesson you can

  • Run the two-chat test on any factual claim.
  • Check an answer against your own book or an official page.
  • Spot the questions that need checking before you ask.
  • Do all this on a phone in under two minutes.

Read first: Why AI makes things up

Someone in your family asks you the last date for a form. You ask an AI app. The answer comes back in two seconds, clean and sure.

Now you have to decide something real. Do you trust it enough to spend a day and a bus fare going to the office?

This lesson gives you four checks. All four fit inside two minutes on a phone.

When a check is worth your two minutes

You cannot check everything. So ask yourself two short questions first.

The first question is about the fact itself. Is it rare, local, or tied to a date?

The model has read a huge amount of text. Common things came up thousands of times there. Your block’s name, one scheme’s exact amount, one person’s date of birth: these may have come up once, or never.

Researchers put a floor under this in 2025. If one fact in five appears only once in the training text, the model gets at least one in five of those wrong. Rare and specific is where it guesses. That is the whole story in why AI makes things up.

Indian questions are weaker still. On a large set of questions from Indian exams, the best model tested got about 74 out of 100 right. That was measured in 2025. It did worst on law, government and culture. There is a full lesson on that gap.

The second question is about you. If this answer is wrong, what does it cost me?

Money, marks, a day of travel, or a missed date. If the honest answer is “nothing”, use it and move on. If it is any of those four, check.

Take two questions from the same afternoon. “What is a fixed deposit?” is common and written about everywhere. Being slightly off costs you nothing.

“What rate does my branch pay on a two year deposit?” is local and tied to a date. It moves real money. The first needs no check at all. The second needs all four.

Your questionRiskWhat to do
What is a percentage?LowUse the answer
Explain this chapter to meLowUse the answer
The last date for a formHighRun all four checks
One scheme’s exact amountHighTrust only the official page
Interest on my loan over 3 yearsHighDo it on your calculator

Check one: ask it again in a fresh chat

Open a new chat. Do not carry on in the old one. Type the same question again, word for word.

Do not mention that you asked before. Do not tell it what the first answer was. Then put the two answers side by side.

Here is a question you can copy:

What is the last date to apply for [name the exam or form]?
Answer in one line.
If you are not sure, write NOT SURE. Do not guess.

Why does this work? Unless you switch on web search, the model does not look anything up. It picks each word by weighing what usually comes next. There is a small amount of chance in that choice, which is why the same question gives different answers.

When the model has truly seen a fact many times, the pull towards it is strong. Chance cannot move it. The answer stays the same in both chats.

When it is guessing, nothing holds it in place. One research team asked a model for a person’s date of birth three separate times. It gave three different dates. All three were wrong.

So the rule is short. Two different answers mean neither one is knowledge. Do not pick the one you like better.

Two matching answers are weaker proof than they feel. The model can be wrong in the same way twice. Matching answers mean “go to check two”, not “this is true”.

One habit spoils this test. Do not put your expected answer inside the question. “The last date is 31 March, right?” pushes the model to agree with you.

Ask it flat, both times. Same words, no hint of what you hope to hear.

Two chats cost you two messages. On a free plan that is a real cost. So spend it on the answers that actually matter.

Try this Ask any AI app for the pincode of your own village. Then ask again in a new chat. Small local facts are exactly where the guessing shows.

Check two: hold it next to something printed

Now find something from outside the chat. Your textbook. A printed notice from an office. An official page whose web address ends in gov.in.

Then paste both into the chat and ask where they disagree:

Here is what you told me:
[paste the AI answer]

Here is what my book says:
[type the lines from your book]

Where do these two disagree?
Which one is right, and how can I check it myself?

This works for one reason. You have handed the model something from outside itself. In one study of self-checking, a model given outside feedback rose from about 76 correct in 100 to about 84 on school maths. That study is from 2024. Left to check itself, the same model fell.

The gap between the two ways of asking is wide. Given the document and told to stick to it, good models slip only a few times in 100. That was the picture in May 2026. Asked instead to recall facts from memory, the same kind of model can be wrong more than 30 times in 100.

Two different jobs, very different odds. And if your book and the answer disagree, your book wins. Printed and dated beats fluent and fast.

So giving it the real document is the strongest habit in this course. Feeding it your own notes and photos beats asking it to remember.

Careful Money, medicine, law and government schemes are different. For these four, an AI answer is never the answer. Confirm at the office, with a doctor, or on the official website.

Sometimes the answer carries a link, or the name of a source. Open it. Every single time.

You are checking two separate things. Does the page open at all? And does that page really say what the AI said it says?

A source name with nothing to open is not a source. If the answer names a book, a rule or a report and gives you no link, treat it as unchecked.

Both fail often. The largest independent test of free AI assistants used news questions. It was published in October 2025. It covered 3,113 questions. Just under half the answers had at least one serious problem. The biggest single problem was sourcing, not raw facts.

Invented names look real on purpose. A 2026 test of five models looked at the software package names they suggested. Roughly 5 to 6 names in every 100 did not exist at all. Web addresses fail in the same way, so a link that looks right may go nowhere.

Turning on web search helps, but it does not close the hole. A hard set of questions was tested in early 2026. On the hardest questions, about 3 in 10 answers with search on still carried a claim the source did not support. Read that beside how it searches and uses tools.

In India this has already cost people money. Lawyers have filed references to judgments that do not exist. Trained people were fooled, because the fake looked exactly like the real thing.

So take the rule that lawyers now teach each other. Never repeat what you have not opened and read.

One more thing about the checking itself. The answer box at the top of a Google search page is also AI. One analysis ran about 4,300 searches. As of February 2026 it found roughly 9 boxes in 100 were wrong. That was better than in October 2025, and Google called the test flawed.

Either way, that box is not your check. Scroll past it and open the real links below it.

On a small data pack, opening pages feels expensive. It is not. One page costs far less than a wasted bus fare, and there are ways to stretch a small pack.

Check four: do every number yourself

Any number that matters goes through your phone calculator. Rupees, marks, interest, land area, days left. All of them.

The model does not see 1234 the way you do. It reads numbers in chunks, not digit by digit. So its slips are not random. They are built in.

It also gets worse as numbers get bigger. In a 2025 test, logical errors rose by up to 14 in 100 as numbers grew. The thinking needed had not changed at all. Wrapping the sum inside a story made it worse again.

So split the job. Let the model set up the calculation. You press the keys.

Write the calculation as separate steps.
Do not give me the final number.
One line per step, like this: 4,500 x 12 = ?
I will do the arithmetic myself.

Then do the last step twice. Once forward, once backward. If the model says 4,500 times 12 is 54,000, divide 54,000 by 12 and see whether 4,500 comes back.

There is more on this in why it gets sums wrong.

“Are you sure?” is not one of the checks

Many videos teach this line. It is the weakest thing you can do.

Here is why. The model has no separate store of facts to look at again. Checking itself only means writing more words, in the same way it wrote the first ones. Nothing new has entered the room.

When models are asked to check their own facts and sums, their scores usually drop. In a 2024 study a leading model fell from about 95 correct in 100 on school maths to about 89 after two rounds. A smaller model fell from about 76 in 100 to about 38 on a general knowledge test.

Asking “how sure are you out of 10” is no better. In a 2026 study the stated confidence barely moved. It stayed high whether the model was right about half the time or three quarters of the time.

If it changes its answer because you pushed, that proves nothing either. It may just be yielding, as correcting a wrong answer shows.

Do not wait for it to admit it does not know. In that study of 3,113 questions, only 17 got a refusal. Saying nothing about doubt is the default, not a sign of knowing.

The working version of “are you sure” is check two. Outside proof, never self-doubt.

Four lines for your notes app

Copy these four lines and keep them. They are the whole lesson.

1. Ask again in a new chat. Two different answers means no answer.
2. Hold it next to my book or an official page.
3. Open every link. Does that page really say it?
4. Every number that matters goes through my calculator.

Use check one on anything rare, local or dated. Use check two before you act on it. Use check three before you repeat it to anyone. Use check four before any money moves.

Nobody runs all four every time. That is fine. Run the two questions from the first section, then pick the checks that fit. Later in this module you will fold all of it into one checking page you keep forever.

There is one more reason a wrong answer can survive every check here. The model wants to agree with you. That is the next lesson.

Do this now

Check the best answer you have got so far

  1. Find the most useful AI answer you have received in this course. Pick one clear fact inside it.
  2. Open a new chat. Type the same question again, word for word. Do not say you asked before.
  3. Put the two answers side by side. Write down whether they agree or not.
  4. Open every link or source name the answer gave you. Check that the page opens and that it really says the thing.
  5. Write three lines in your notes app: what matched, what did not, and what you still need to check at an office or in a book.

Remember this much

  • Check when the fact is rare, local or tied to a date, and when being wrong costs money, marks or a day of travel.
  • Ask the same question in a second fresh chat. Two different answers mean neither one is knowledge.
  • Two matching answers are a hint, not proof. The model can be wrong the same way twice.
  • Outside proof is the only real proof. Your book, a printed notice, or an official page.
  • Open every link before you repeat what it says. Real-looking sources are the most convincing mistakes.
  • Are you sure is not a check. Models usually get worse at facts when they check themselves.

Questions people ask

How do I check if a ChatGPT answer is correct?

Ask the same question again in a brand new chat and compare the two answers. If they differ, neither is knowledge. Then hold the answer next to your textbook or an official page whose address ends in gov.in. Open every link the answer gave you before you repeat it to anyone.

Is it enough to ask the AI are you sure?

No. When models are asked to check their own facts and sums, their scores usually drop. In a 2024 study a leading model fell from about 95 correct in 100 on school maths to about 89 after two rounds of self-checking. Outside proof works. Self-doubt does not.

Why do I get a different answer when I ask the same question twice?

The model picks each word by weighing what usually comes next, and there is a small amount of chance in that choice. When it truly knows a fact, chance cannot move it and the answer stays the same. When it is guessing, the guess changes. That is what makes the two-chat test useful.

Is the AI answer at the top of Google search reliable?

It is AI too, and it is sometimes wrong. One analysis of about 4,300 searches found roughly 9 boxes in 100 were wrong as of February 2026, which was better than in October 2025. Google disagreed with the method used. For anything that matters, scroll past the box and open the real links.

Can I trust an AI answer about a government scheme or a last date?

Not on its own. Money, medicine, law and government schemes are the four areas where an AI answer is never the answer. Use it to understand the words and to make your list of questions. Then confirm at the office, or on the official website.

Prices, free limits and app screens change often. The facts in this lesson were checked on 9 August 2026. If what you see on your phone looks different, trust your phone and read the idea, not the exact button name.

See all 64 lessons