Home / Journal / Deep Dive

How AI Decides What Is True About a Person

Deep Dive2026-07-1212 min read
The machinery

No engine checks facts about you the way a careful journalist would. Models weigh corroboration: how often a claim appears, in which sources, and how well it agrees with what they already hold. The version of you an AI repeats as truth is really a vote among documents. The good news hiding inside that unsettling sentence: you can influence the electorate.

Ask a chatbot about yourself and it answers in the flat, assured tone of a reference book. It never says "I believe" or "sources disagree" unless pushed. Where does that certainty come from? Not from checking. From counting.

This matters more than it used to, because the machine's version of you is now read aloud to buyers, hiring committees, journalists and conference organisers before you ever enter the room. Understanding how that version gets assembled, which is to say how a statistical system decides what is true about a person, is the first step to having any say in it.

Where does an AI's version of you come from?

Two doors, and it helps to keep them separate. The first is training: everything the model absorbed before its knowledge cutoff, compressed into parameters. Whatever the web said about you during training, repeated often enough to leave an imprint, is what the model "knows" when it answers from memory. The second is retrieval: many AI answers now begin with a live search, and the model reads a handful of current pages before composing its reply. OpenAI even runs distinct crawlers for the two jobs, GPTBot for training data and OAI-SearchBot for its search index, with ChatGPT-User fetching pages during live browsing (OpenAI's bot documentation lists all three; Search Engine Land maintains a broader guide to AI crawlers).

The two doors have different clocks. Retrieval can pick up a correction you published on Tuesday. Training bakes in whatever the web said months or years ago and holds it until the next model. Which is why a fixed error sometimes stays fixed in one product and resurrected in another: you corrected the live web, but the memory still carries the old vote.

How does AI verify information?

Bluntly: it doesn't, not in any sense your fact-checking teacher would recognise. There is no module that rings your university to confirm the degree. What the system has instead are three proxies that usually correlate with truth, and sometimes spectacularly don't.

Frequency. A claim that appears in many independent documents is weighted as more likely true than one appearing once. Provenance. A claim in a source the system treats as reliable, an established publication, a curated database, an institutional page, counts more than the same claim on an anonymous blog. Coherence. A claim that fits neatly with everything else attached to your entity is accepted smoothly; one that clashes triggers hedging or omission.

Notice what all three have in common: they measure agreement, not accuracy. The machine's working definition of truth about you is "what trusted documents keep saying, consistently." Most of the time this approximates reality rather well, which is exactly what makes the failures so hard to spot.

the core mechanism

AI treats corroboration as verification. A fact about you becomes "true" when enough independently trusted documents repeat it without contradiction. Truth, to the machine, is a quorum.

Why can consensus beat accuracy?

Because errors replicate. A journalist misstates your job title in one profile; three aggregator sites syndicate it; a conference bio copies the aggregator; a database scrapes the conference. To a system counting votes, that single mistake now looks like four independent confirmations, while your correct bio, published once on your own site, casts one vote. The wrong version wins the election, and the model repeats it with perfect confidence.

This is the anatomy behind most of the horror stories readers send us, the kind we dissected in What To Do When AI Gets Facts Wrong About You. It is also why the fix is never to argue with the chatbot. The chatbot is downstream. You have to fix the sources casting the votes, and there is a related, stranger failure worth knowing about: when no quorum exists at all, models sometimes fabricate one, inventing degrees and employers that fit the pattern of people like you. That failure mode gets its own treatment in When AI Invents Your Credentials.

Which sources get believed first?

Not all votes weigh the same. Reconstructing the hierarchy from how engines behave, and from what Google publishes about its own reliability standards in its guidance on helpful, people-first content and E-E-A-T, gives you roughly this ladder:

The believability ladder
Source typeWhat machines take from itWhy it ranks where it does
Curated databasesCore identity facts: name, roles, dates, affiliationsWikipedia and Wikidata are structured, cross-referenced and policed, the closest thing to a machine's notary
Established mediaEvents, achievements, quotes attributed to youEditorial process implies a human already checked, and other sources cite them, multiplying the vote
Institutional pagesCredentials, memberships, employmentUniversities, regulators and professional bodies rarely publish flattery
Your own siteCanonical spellings, titles, bio, structured dataAuthoritative for identity details, discounted for self-praise; the machine knows who wrote it
Forums and socialSentiment, reputation texture, recencyHigh volume, low individual trust; matters in aggregate, not per post

Two practical readings of that table. First, the top rows are why we bang on about curated databases in Wikipedia, Wikidata and Why Machines Trust Them: a fact lodged there propagates downward through everything else. Second, your own website sits in a peculiar position: it cannot make the machine believe you are brilliant, but it is the reference implementation for the boring facts, how your name is spelt, what you actually do, which profiles are really yours. Getting those boring facts consistent everywhere is most of the game.

What happens when sources disagree about you?

Three outcomes, in rough order of likelihood. The machine goes with the majority, which is fine when the majority is right and maddening when it isn't. Or it hedges, blending both versions into something vaguely wrong. Or, and this is the quiet killer for professionals, it omits you entirely, because a system asked to recommend someone trustworthy will skip the candidate whose record contradicts itself and name the rival whose record is boringly coherent instead.

There is a fourth case, gentler but more common: the machine believes a version of you that has simply expired. You changed firms, retrained, moved countries, narrowed your specialism, and nobody told the record. The old truth was never wrong, so no source contradicts it; it just quietly outnumbers the new one. If your career has had chapters, assume the engines are still reading an earlier chapter until you have republished the current one widely enough for it to win the vote.

Disagreement also arises when the machine cannot tell whether two records describe one person or two. An old profile under a nickname, a different title on every platform, a namesake in another industry: each fragment splits the vote. That resolution problem, and the consolidation that fixes it, is the subject of our piece on being an entity, not an ego.

Does the truth about you expire?

Faster than you would think. Semrush's AI Visibility Index found that 40 to 60 percent of the sources cited in AI answers rotate month over month (via Similarweb's generative AI statistics roundup). The documents casting votes about you this month are substantially different from last month's electorate. A record that was accurate in 2023 and untouched since is not stable; it is decaying, because newer, possibly sloppier documents are constantly joining the pool while your correct ones age out of favour.

The uncomfortable conclusion: truth about a person, as far as machines are concerned, is not discovered once. It is maintained, or it drifts.

The corroboration audit: six steps

Here is the maintenance routine we run, reduced to what you can do in an afternoon:

  1. Collect the machine's current version. Ask ChatGPT, Perplexity and Gemini who you are and what you are known for. Save the answers verbatim, with dates.
  2. Trace every claim to its voters. For each fact, right or wrong, find which pages assert it. Retrieval-based engines often cite theirs; for the rest, search the exact phrasing.
  3. Fix errors at the highest-ranking source first. A correction on an established publication or a curated database outweighs ten corrections on minor profiles.
  4. Publish one canonical record. A bio page on your own domain with consistent name, title and history, marked up with structured data, linked from every profile you control.
  5. Retire the contradictions. Old bios, stale titles, abandoned profiles: update or delete. Every contradiction is a vote against your own coherence.
  6. Re-test quarterly. Because of citation rotation, a clean answer today is not a clean answer in October. Put it in the calendar.

None of this is glamorous, and that is rather the point. The people machines describe accurately are rarely the loudest; they are the ones whose record agrees with itself everywhere the machine looks. If you would rather have the audit run for you, complete with the source-by-source trace, that is what our check-up does.

What happens when a bio still says the old thing?

The abstract version of this problem is easy to state, and the concrete version is where people actually get stuck, so it helps to walk through one. Imagine a specialist named Priya whose law firm bio has listed her as "Associate" since 2021. She made partner two years ago, and since then the firm's own updated site calls her "Partner," a legal directory entry calls her "Partner," and three separate conference programs list her the same way. The old "Associate" bio never got taken down. It just sits there, technically accurate once, now simply out of date.

Ask an AI assistant what Priya does and you are watching a small election happen in real time. In the months right after the promotion, the old bio was still the only vote on record, so an assistant answering from memory alone would repeat "Associate" with total confidence, because one vote was all there was. As the newer mentions accumulated, independent, in different formats, on sites that plainly did not copy from each other, the newer fact crossed from a single claim into a corroborated one. Once that threshold is crossed, retrieval-based answers tend to flip first, because they are reading current pages at the moment of the question. Answers drawn purely from training data lag behind, sometimes for months, until a new model absorbs the updated web.

The practical lesson sits in that lag. If you have had a real change, a promotion, a new firm, a rebrand, do not expect the correction to be instant just because it is now technically true. It becomes true to the machine once enough independent sources say so and the old version stops outnumbering the new one. Until then, both answers are plausible outputs of the same honest mechanism, and that is worth knowing before you assume the model is simply being sloppy.

the corroboration climb
  1. Step 1: Single mention seen. One page states the fact. The system notes it but treats it as unconfirmed.
  2. Step 2: Repeated across sources. Independent sites that clearly did not copy from one another start saying the same thing.
  3. Step 3: Treated as reasonably reliable. Once independent repetition crosses a rough threshold, the system starts using the fact as a working assumption.
  4. Step 4: Surfaced in an answer. The fact graduates from background assumption to the sentence an assistant actually says when someone asks.

How a single claim about you climbs from an unconfirmed mention to something an AI will confidently repeat.

How do you correct the record, if asking the chatbot doesn't work?

The instinct, understandably, is to go straight to the tool people notice: type a correction into the chat window itself, or fill out a feedback form on the provider's site. Both are worth doing, and both are close to useless on their own. A correction typed into a chat only affects that one conversation. It does not touch the underlying sources, so the next person who asks starts from the same vote count you just tried to argue with.

Correcting at the source means finding the pages actually casting the incorrect votes and fixing them there, in roughly this order of priority. Start with whatever ranks highest on the believability ladder above: a curated database entry, if one exists, then an established publication that ran the error, then institutional pages, then your own site. A single correction on a high-trust page is worth more than a dozen corrections on minor ones, because the system counts corroboration, not effort. If you would rather have someone else run that source-by-source trace for you, that is what our check-up covers, and it pairs naturally with the harder, related question of whether you can ever make AI forget an old fact entirely, which is a different problem from simply updating one.

The last step is patience dressed up as a plan. Once the source pages are fixed, the new version has to accumulate its own corroboration before it can outvote the old one, and that takes real time, not a support ticket. How long that actually takes depends on how many places pick up the correction and how quickly engines re-crawl them, but the direction is always the same: fix the pages, not the chatbot, and let the count catch up.

The moral of the machinery

It is tempting to find all this sinister: truth about you decided by a vote you never see, among documents you didn't write. But consider the alternative reading. For the first time, the mechanism of reputation is legible. We know what the machine counts, where it looks, and how it breaks ties. Reputation among humans was always a whisper network; this one publishes its rules. The professionals who thrive in the next decade will simply be the ones who read them.

Questions people ask

How does AI verify information about a person? +
It doesn't verify in the human sense. Models weigh corroboration: how often a claim appears, in which sources, and how well it agrees with what they already hold. Agreement across independently trusted documents is treated as truth.
Why does AI repeat wrong facts about me? +
Usually because an error was published somewhere reasonably trusted, then copied. Once several sources carry the same mistake, it looks like consensus, and consensus is what the machine mistakes for accuracy.
Can I change what AI says is true about me? +
Yes, but slowly and at the source. Correct the pages engines actually read, publish one canonical bio with structured data, and earn fresh corroboration from independent sites. Retrieval-based answers update first; trained models follow later.
What happens when two sources disagree, like an old job title versus a newer one? +
The older claim keeps being repeated until the newer one is corroborated by enough independent sources to outnumber it. Retrieval-based answers tend to catch the change first, since they read current pages; answers from training data lag until a new model absorbs the updated web.
Does asking a chatbot to fix a wrong answer actually correct the record? +
No. Typing a correction into a chat window only affects that one conversation. The underlying source pages still carry the error, so the next person who asks starts from the same vote count. Fix the source pages, starting with the highest-ranking ones, instead.
How long does it take for a correction to actually take hold? +
There is no fixed timeline. It depends on how many independent places pick up the correction and how quickly engines re-crawl them. Retrieval-based answers can update within weeks; answers drawn from training data can lag for months until a new model is released.

Curious what AI says about you?

Start with a check-up. We'll show you the exact words the engines return about your name, then map the fastest signal to move.

Say my name →