Home / Journal / Explainer

How Does ChatGPT Know Things? Training vs Looking Things Up

Explainer2026-07-278 min read
Short answer

ChatGPT knows things in two genuinely different ways. It has trained knowledge, patterns absorbed from a huge pile of text up to a fixed date called the training cutoff, and some versions also have live search, the ability to actually browse the current web while answering your question. Trained knowledge is frozen in time. Live search is current but depends entirely on what it can find and trust in the moment. Most confusing "why doesn't it know that" moments come down to mixing these two up.

People often assume ChatGPT works like a search engine, typing in a question and getting back whatever's currently true on the internet. It's actually running two different processes that can behave very differently, and knowing which one is active in a given answer explains a lot of otherwise strange behavior.

Trained knowledge: a frozen snapshot

When a model is trained, it's shown an enormous amount of text gathered up to a certain point in time, and everything it learns comes from that fixed collection. Once training finishes, that knowledge doesn't update on its own. This fixed point is called the training cutoff, the date after which the model simply has not read anything, because nothing after that date was part of what it was trained on. Ask a model without live search about something that happened after its cutoff, and it genuinely has no way to know, the same way you couldn't know about a news story you never read.

Live search: current, but conditional

Many modern AI tools, including versions of ChatGPT, Gemini, and Perplexity, are also connected to the live web, meaning they can send out an actual search in the moment you ask a question, read a handful of current pages, and use that fresh information to shape the answer. This gets around the training cutoff problem for things that have happened recently, but it comes with its own condition: the tool can only use what it can actually find, what's indexed by the search system it relies on, and what it judges to be a reasonably trustworthy source in that moment. If your recent material hasn't been indexed yet, or isn't easily found, live search won't magically surface it either.

Trained knowledgeLive search
When it was learnedFixed, up to the training cutoffRight now, at the moment of the question
Can it update?Only when the model is retrainedYes, each time it searches
Depends onWhat was in the original training dataWhat's currently indexed and trusted
Typical useGeneral knowledge, reasoning, established factsRecent events, current pricing, fresh mentions

Figure: the two knowledge sources behave differently, and most answers blend the two.

Why this explains so many confusing moments

This split explains a lot of the head-scratching people run into with these tools. It explains why a model can be confidently wrong about something that changed recently, it simply hasn't read the update yet. It explains why the exact same question can get a different answer depending on whether a tool has live search turned on. And it explains why publishing something new about yourself doesn't instantly show up everywhere, since a training-based answer won't reflect it until the next retraining, and even a live-search answer needs your new material to be indexed and trusted first.

What this means for you specifically

If you want a live-search-enabled tool to find and mention something you just published, the practical requirement is that it needs to be genuinely findable, meaning indexed by the search systems these tools rely on, and clear enough that the tool trusts it as a real, attributable source. If you're hoping to shape what a model says from its trained knowledge, that's a slower game, since it only updates when the model itself gets retrained, which isn't something on any predictable public schedule you can plan around.

A simple way to test which one you're dealing with

Ask about something you know happened very recently, in the last few weeks, and see whether the tool knows about it at all, and whether it mentions searching or citing a source when it answers. If it correctly describes the recent event and shows signs of having pulled from a live source, you're seeing live search at work. If it either doesn't know about it or gives an outdated answer, you're likely seeing pure trained knowledge with no live search involved in that particular response.

Why the same question can get different answers on different tools

Even setting the training-versus-search question aside, different AI tools were trained on different collections of text, with different cutoffs, and different companies have made different choices about how heavily to weigh live search versus trained knowledge in a given answer. None of this makes one tool simply "more correct" than another in some universal sense, they're just built differently, drawing from different combinations of frozen and fresh information. If you want a direct comparison of how the major tools actually differ in practice, Perplexity versus ChatGPT versus Gemini lays it out side by side.

Why publishing consistently still matters either way

Whether a particular answer comes from trained knowledge or live search, the underlying requirement is the same: there needs to be clear, substantial, attributable material about you somewhere for either process to draw from. Trained knowledge needs it to have existed before the cutoff. Live search needs it to be findable and trustworthy right now. Publishing consistently over time gives you a shot at both, rather than betting everything on one mechanism catching you at the right moment.

Why this matters most for breaking news and very recent changes

The training-versus-search distinction becomes most obvious around anything that changed recently, a new role you took on, a business that recently moved locations, a fact that used to be true and no longer is. A tool relying purely on trained knowledge will confidently repeat the old, outdated version, since that's genuinely what it learned. A tool with live search has a real chance at getting it right, but only if the updated information is already indexed and appears trustworthy enough to surface. This is exactly why keeping your own published material current matters just as much as publishing it in the first place.

How this shapes where you should focus your effort

If you're hoping a live-search-enabled tool picks up something quickly, prioritize making sure it's easy to find and clearly credible, a well-structured page, indexed properly, is far more useful here than a hard-to-reach one. If you're thinking about longer-term trained knowledge, accept that it moves on its own slower schedule and focus instead on building a body of consistent, substantial material over time, so that whenever the next training snapshot is taken, there's more for it to learn from. Where AI looks before recommending you goes further into the specific sources these tools tend to lean on.

A simple way to remember the difference going forward

If it helps, picture trained knowledge as a well-read friend's memory of every book they finished before a certain date, fixed and unchanging until they read something new. Picture live search as that same friend stepping out of the room for a moment to check today's newspaper before answering you. Both are useful, both have real limits, and knowing which one is answering your specific question tells you exactly how much to trust its freshness.

Questions people ask

Does ChatGPT always search the internet when I ask it something? +
Not always. It depends on the specific version and settings being used. Some responses rely purely on trained knowledge, while others actively search the live web, and the tool sometimes indicates which one it's doing.
What exactly is a training cutoff? +
It's the date up to which the text used to train a model was gathered. Anything that happened or was published after that date simply wasn't part of what the model learned from, unless it also has live search ability to find it separately.
Why did an AI tool know about something very recent? +
That's almost certainly a sign it used live search rather than relying purely on trained knowledge, since a training cutoff by definition can't include anything that happened after it.
Can I make an AI tool ignore its training and only use live search? +
Some tools let you nudge this by explicitly asking them to search or check current sources, though the exact behavior depends on the specific tool and whether that feature is available and enabled.
Does a newer AI model always know more recent things? +
Generally yes, since a newer model typically has a later training cutoff, though the exact date varies by model and company, so it's worth checking rather than assuming.
If I publish something today, when might an AI tool know about it? +
A live-search-enabled tool could potentially find it once it's indexed, which can take days to weeks. A purely trained model won't reflect it until that model's next training cutoff and release, which can take considerably longer.

Curious what AI says about you?

Start with a check-up. We'll show you the exact words the engines return about your name, then map the fastest signal to move.

Say my name →