Lesson 16 of 64 Module 3: What is really happening inside

Why AI forgets what you said: the context window

8 min read Free, no sign-up 9 August 2026

After this lesson you can

  • Explain the difference between what it learned and what it can see.
  • Say what happens when a chat gets too long.
  • Keep important instructions where they will be read.
  • Decide what to paste and what to leave out.

Read first: It is guessing the next word

In your first message you said you have passed class 10 and are preparing for a railway exam. Thirty messages later, the same chat hands you a study plan for a bank job. You changed nothing, and your phone is not at fault.

The chat ran out of room. Nothing on the screen told you, so this lesson tells you instead.

What it learned, and what it can see

Two different things sit behind every answer you get.

The first is what the model learned before you ever typed. It read a huge amount of text during training, and that reading is now fixed inside it. Nothing you type changes it. How a model is made covers that part.

The second is what it can look at right now. Picture a person sitting at a desk. Everything they studied is inside their head. In front of them is a table. Only what is on that table can be read at this moment.

That table has a size. It is called the context window.

Your messages go on the table. So do its own replies. So does every file, photo and screenshot you send. So does a hidden note the company adds before your message. When the table is full, something has to leave it.

This one picture kills the most common wrong idea. Pasting a chapter does not teach the model. It puts the chapter on the table for this chat only. Open a new chat tomorrow and that chapter is gone.

Keep these two things apart in your head and most of the confusion goes away.

QuestionWhat it learned in trainingWhat is on the table
Where it came fromText read before you typedThis chat, right now
Can you change itNoYes, with every message
How long it lastsUntil the company builds a new modelUntil this chat ends or fills up
Your notes and photosThey never go hereThey go here, for this chat only

It reads the whole chat again, every time

Here is the part that surprises people. Nothing is running between your messages. The model is not sitting there thinking about you while you type your next line.

When you press send, your whole conversation is placed on the table again from the beginning. Your first question, its first answer, all of it, and then your new line. It reads the pile, writes one reply, and holds on to nothing.

So it feels like it remembers you. That feeling comes from the pile being re-read, not from a mind that stays awake.

One more thing lands on that table, and few people think of it. When the app works through a problem in steps, that rough work is text as well. It takes up room like everything else. What thinking mode really does explains when that room is worth spending.

Each turn puts a bigger pile in front of it. A long chat means more to read before every single reply. If your data pack is small, using AI without finishing your data has practical ways to spend less.

What happens when the table fills up

Three things can happen. None of them announce themselves.

In a chat app, the oldest part of the conversation usually slides off the table. First in, first out, like a queue at a bank counter. Your early messages are simply not there any more.

Or the app squashes the old part into a short summary and keeps that instead. The gist survives. Your exact words and small details do not.

In the tools programmers use, the request is refused with a plain error saying the text is too long. You will not see that in a chat app. You just get an answer that has quietly lost the start of your story.

Notice what is missing in a chat app. There is no warning, no red line, no message. The chat looks exactly as it did on day one.

You can still spot it afterwards. It asks for your class or your district again, although you gave both at the start. It repeats a point it already made. It answers the question you asked two turns ago. Read all of that as a full table, not as a bad mood or a slow phone.

Worth knowing Photos take up a lot of room on that table. One clear photo of a page can use as much space as a page of typed words. Send the one page you actually need.

A bigger table is not a better reader

Apps now advertise very large memory sizes. That sounds like the model can read a whole book properly. It cannot, and you deserve the honest version.

In 2025 a team tested 18 leading AI models on long inputs. Accuracy dropped as the input grew, well before the stated limit was reached. Adding text that looked related but was not made the answers worse again.

There is a reason inside the machine. Every piece of your text is compared against every other piece. Double the text and the work goes up four times. Its attention is a budget, and a long chat spreads that budget thin.

The effect is one you have probably felt. The start and the end of a long chat get used. The middle gets neglected. The detail you gave on message fifteen is the one it drops.

So do not hand it thirty pages and hope. Give it one part, take the answer, then start again with the next part. Small pieces, one at a time, beat one big pile.

Three habits that fix most of this

Rule one. Paste the paragraph, not the whole book.

If your doubt is about one method in a chapter, send that method alone. A short clean paste beats a long messy one every time. Giving it your own book, notes or photo shows you how to pick the right piece.

Rule two. Put your main instruction first, then say it again at the end.

In a long message the middle is the weakest place. So state what you want, then paste the material, then repeat the instruction below it.

Answer in simple Hindi, in under 150 words.

Here is the passage from my book:
[paste your passage here]

Now answer, in simple Hindi, in under 150 words.
Use only what is in the passage above.
If the passage does not say it, write NOT SURE.

A good reply to that stays inside your passage and stays short. If it wanders off into general knowledge, your passage was too long. Cut it down to the part you really need and send again.

Rule three. One chat, one job.

You met this rule in module 2. Now you know the machinery behind it. A new topic means a new chat, because every old topic is still sitting on the table taking up room. When a chat has gone stale, move your work into a fresh one instead of arguing with it.

There is one rescue move for a chat you cannot leave. Send a short message that holds only the facts that matter, written out again in your own words. Fresh text at the bottom of the pile does get read. A detail buried twenty messages up may not.

Hindi fills the table faster

The model does not count your text in words. It cuts text into small pieces first. Each piece is called a token. The lesson on how it is guessing the next word explains where those pieces come from.

Devanagari gets cut into many more pieces than English for the same meaning. One 2026 study measured Hindi needing about four times as many pieces as the same content in English, on one widely used older system. Newer systems have cut that gap a lot, and it keeps improving.

Here is what that means for you today. A long Hindi document fills the table several times faster than the same document in English. For a short question, the difference is too small to worry about.

So do not give up Hindi. For a long paste, use English if you can read it, and ask for the answer in Hindi. Using AI in Hindi has the full method.

When the answers still go strange

You now hold two of the big reasons an answer goes wrong. The model is guessing the next piece of text. The table it reads from has a size, and it drops things without telling you.

There is a third reason, and it is stranger than both. Ask the same question twice, in the same words, and you can get two different answers. That is not a fault and it is not the app playing with you. The next lesson, on why the same question gives different answers, shows what is really going on.

Do this now

Whole book against one paragraph

  1. Find something long on your phone. Two or three pages of a chapter or an article is enough.
  2. Open a new chat. Paste the whole thing in, then ask one question about a small detail from the middle of it.
  3. Read the answer. Check it against the text yourself and mark it right, wrong or vague.
  4. Open a second fresh chat. This time paste only the one paragraph that holds the answer.
  5. Ask the same question in the same words. Compare the two answers and write one line about what changed.

Remember this much

  • What it learned in training is fixed. What it can see is only this chat.
  • There is no memory between messages. The whole chat is read again every turn.
  • When the space runs out, the oldest part drops off or becomes a short summary.
  • Nothing on the screen warns you when that happens.
  • A bigger memory is not a better reader. The middle of a long chat gets neglected.
  • Paste the paragraph, not the whole book, and repeat your main instruction at the end.

Questions people ask

What is a context window in AI?

It is the amount of text the model can look at while it writes one answer. Your messages, its replies, your files and your photos all sit inside it. It is separate from what the model learned during training. Think of it as the table in front of it, not the knowledge in its head.

Why does ChatGPT forget what I said earlier?

Because your conversation grew bigger than that working space. Chat apps usually drop the oldest messages or replace them with a short summary so the chat can carry on. Your early details are then gone, and nothing on the screen tells you. Starting a fresh chat with a short summary of your own fixes it.

Does uploading a file or photo teach the AI?

No. The file only enters the working space for that one conversation. The model itself does not change at all. Open a new chat tomorrow and the file is not there. If you need it again, send it again.

The app remembers my name, so does it have memory?

Some apps keep a small notebook of text about you and read it back before answering. That is a saved note, not the model learning. You can usually see and delete these notes in the settings. On a shared phone, check them, because the next person using that account may see them too.

Does Hindi use up the AI memory faster than English?

Yes, for long text. Devanagari gets cut into many more small pieces than English for the same meaning, so a long Hindi document fills the space faster. For a short question the difference does not matter. For a long paste, English fills less space, and you can still ask for the answer in Hindi.

Prices, free limits and app screens change often. The facts in this lesson were checked on 9 August 2026. If what you see on your phone looks different, trust your phone and read the idea, not the exact button name.

See all 64 lessons