Lesson 15 of 64 Module 3: What is really happening inside

How an AI model is trained, in three plain steps

10 min read Free, no sign-up 9 August 2026

After this lesson you can

  • Name the three stages that make an AI model.
  • Say where its knowledge comes from, and where it does not.
  • Explain why it knows little about recent months.
  • Show a friend why chatting with it does not teach it.

Read first: It is guessing the next word

Yesterday you corrected the app. You told it the right name of your block office, and it thanked you. Today you open it again and the mistake is back.

A friend told you it learns from every chat. It does not. This lesson shows you what actually made it, and why that matters for the way you use it.

Stage one: it reads, and learns how writing continues

Nobody sat down and typed the answers into this thing. That is the first surprise.

It is made in three stages. First it reads. Then it learns to answer. Then people correct the kind of answer it gives.

Stage one is reading, on a scale that is hard to picture. The makers feed it an enormous amount of writing. Web pages, books, manuals, news, question papers, conversations between people.

This stage runs for months, inside a building full of special computers. It costs a huge amount of money and electricity. That is why only a handful of organisations in the world make a model this size.

From all that reading it learns one skill. It learns how writing usually continues. The lesson on guessing the next word takes that skill apart in full.

At the end of stage one you have something strange. It is called a base model. It continues text. It does not answer you.

Give a base model this line, and it may continue like this:

how do I write a leave letter
how do I write a bank application
how do I write a complaint to the block office

That is not an answer. That is a web page continuing. On a real page, one question is often followed by more questions, so that is what it produced.

Stage two: people show it how to answer

Stage two teaches a new habit. Do the job asked, then stop.

People write example conversations by hand. A question, then a good answer to that question. A very large number of them, across many topics.

The model is trained on those examples. It picks up the shape of an assistant. Someone asks, you answer the thing asked, you stop.

This is why the app replies to you instead of writing three more questions. That behaviour was put in on purpose, in stage two.

Underneath, it is still the same model from stage one. Same reading. Only the habit is new.

You never meet a base model in a normal app. The app hands you the trained assistant. Still, it helps to know that a plainer thing sits underneath, because that plainer thing is doing the work.

Stage three: people mark which answer is better

Stage three is correction. It is the stage almost nobody has heard of, and it shapes what you see most.

The model writes two answers to the same question. A person reads both and marks which one is better. This is done a very large number of times.

Training then pushes the model towards the kind of answer people preferred. Not towards one exact answer. Towards a type of answer.

Notice what that does not give it. It does not give it one fixed reply to hold on to. It learned a preferred style. So do not expect the same words twice, and see why the same question gives different answers for the rest of that story.

Nearly everything you notice about its manner comes from here. The politeness. The habit of giving you headings and numbered steps. The refusals when you ask for something harmful.

Some of this marking is now done against a written rulebook instead of one person’s taste. One company, Anthropic, published the rulebook it uses for its model in January 2026, and made it free for anyone to read.

One thing matters more than the details. All of this happened before you opened the app. Nobody is sitting there reading your chat and marking it now.

StageWhat happensWhat you end up with
One. ReadingIt reads a huge amount of writingSomething that continues text
Two. ExamplesPeople write good answers by handSomething that answers, then stops
Three. MarkingPeople mark which answer is betterIts manner, its format, its refusals

Why it knows nothing about last month

Stage one ended on a date. The reading stopped there. That date is called the knowledge cut-off.

Think of a newspaper that stopped arriving at your door. Everything after that day is simply missing.

There is a second part that most people miss. The months just before the cut-off are thin as well. Writing about any event keeps appearing for months afterwards. If the reading stopped soon after the event, it caught only a little of that writing.

So the newest facts it holds are also its weakest facts. Not only the ones after the date.

There is a small oddity worth knowing. The app usually knows today’s date. That is not memory. The company writes the date into a hidden note, and that note is added in front of your message every single time.

The date is different for every model, and it moves with every new version. Do not memorise one. Ask the app instead, and ask whether it can look things up now:

Can you search the internet right now?
Answer yes or no in one line.
Then tell me how I can see on screen that you searched.

If it can search, it can go past its cut-off. That is the real fix for recent news, and how it searches, sees and uses tools explains what is happening when it does.

For an exam date, a result, a scheme rule or a price, do not stop at the app. Open the official site and read it there.

Talking to it does not teach it

Read this line twice. Nothing you type changes the model.

Training happened once, in a building far from you, before you opened the app. After that, the model is frozen. Your chat is not a lesson for it.

There is no running mind between your messages either. Each time you press send, your whole chat is sent again from the start. The model reads all of it fresh, then writes the next reply.

Correcting it inside a chat does work. It works for that one chat, and only until the chat ends. Tomorrow, in a fresh chat, the same mistake returns. The lesson on correcting it when the answer is wrong shows what to say while you are still in the chat.

Try this Tell it to answer everything in exactly three short lines. It will obey for the rest of that chat. Now open a brand new chat and ask anything. The three-line rule is gone. That is this whole lesson in one minute.

Some apps now offer a memory feature. It looks like learning. It is not learning.

What it does is keep notes about you in a file. Before your next answer, those notes are quietly added to the chat. The model reads them the same way it reads anything you typed.

The model itself does not change. The notes are usually plain text you can read, edit and delete. Look in the app’s settings for a word like memory or personalisation.

Careful Phones get shared at home. Anyone who opens that app on that phone can read the memory notes, and the model will use them in their answers too. If you share a phone, keep memory off, or check what is saved in it.

The lesson on what happens to what you type goes further into this.

What it read decides what it is good at

Go back to stage one for a moment. Whatever it read a lot of, it is strong at. Whatever it read little of, it is shaky at.

It read far more English than Hindi. It read far more about America than about your district. That is not an opinion about India. It is a count of pages.

You can feel this yourself. Ask about a national exam and the answer is usually solid. Ask about your block office, a local college, or a scheme your neighbour applied for, and it turns vague or confident and wrong.

Try this on a place you know well:

Name three things you know about my district.
My district is <write your district> in <write your state>.
After each one, say how sure you are, and say if you are guessing.

Read its answer beside what you already know. Why it is weaker on Indian questions works through this properly in module 4.

There is a sensible way to work with this. Let it explain how a thing works in general. Get the local detail from the office itself, from the notice board, or from a person who has already done it.

This is changing a little. In February 2026, three Indian models were shown in Delhi, built to cover the 22 scheduled Indian languages. They are much smaller than the world’s biggest models. Their strength is Indian languages, not being the cleverest model on earth.

Where this leaves you

One honest thing is left. Even the people who build these models cannot fully explain what goes on inside them.

They choose the reading. They choose the marking. They do not write the rules the model ends up following. One of the people who builds them puts it this way: these systems are grown more than they are built.

That will make sense to anyone who has grown a crop. You choose the seed, the soil and the water. You do not decide the shape of each leaf.

So the right position is checked trust, not blind trust. Use it freely, then check the parts that would cost you something if they were wrong. The two-chat test and three other checks gives you four checks that take under two minutes.

You now know the model is frozen and learns nothing from you. So where does the thing you typed five minutes ago actually live? It sits on a small table in front of the model, and that table fills up. Why it forgets what you said is next.

Do this now

Find the edge of what it knows

  1. Open any AI app you can reach on your phone.
  2. Ask it: what is your knowledge cut-off date? Answer in one line.
  3. Now ask it about something that happened last month, near you or in the news.
  4. Watch which of three things it does: says it does not know, guesses, or searches.
  5. If it searched, look for the small line on screen that names the source. Open that source.

Remember this much

  • Three stages make it: reading, then example answers, then human marking.
  • Nobody typed the answers in. It learned how writing continues.
  • Its reading stopped on a date, and the months before that date are thin.
  • Nothing you type changes the model. Training finished before you opened the app.
  • Memory features keep notes in a file. The model itself does not change.
  • It read far more English than Hindi, so it is weaker close to home.

Questions people ask

How are AI models trained?

In three stages. First it is shown an enormous amount of writing and learns how writing continues. Then people write good example answers, and it learns to answer and stop. Then people mark which of two answers is better, and training pushes it towards the preferred kind.

Does ChatGPT learn from my chats?

Your chat does not change the model. Training happened once, before you opened the app, and the model is frozen after that. A correction you make lasts only inside that one chat. Whether a company stores your chats for future training is a separate question, and you check that on the app's own settings page.

What is a knowledge cut-off date?

It is the date the training reading stopped. After that date the model knows very little. The months just before it are also thin, because writing about an event keeps appearing for months. So recent facts are its weakest facts.

What is RLHF in simple words?

People compare two answers from the model and mark which one is better. Training then moves the model towards that kind of answer. This is where its politeness, its refusals and its habit of using headings and steps come from.

If AI has memory, is it learning about me?

No. A memory feature saves notes about you in a file. Before your next answer, those notes are added to the chat as text. The model reads them and is otherwise unchanged. On a shared phone, anyone using that account can see and use those notes.

Prices, free limits and app screens change often. The facts in this lesson were checked on 9 August 2026. If what you see on your phone looks different, trust your phone and read the idea, not the exact button name.

See all 64 lessons