Back to blog Video as a Source: What AI Actually Reads on YouTube

Video as a Source: What AI Actually Reads on YouTube

A model never watches your video. It reads the text around it. Which means a video without a clean transcript is, to a machine, almost mute.

Most marketing teams file video under "brand". It lives with the logo and the photography, it gets measured in views and watch time, and nobody in the room ever asks the question that now matters: can a language model read it?

It can, or rather it can read something that stands in for it. Video has quietly become one of the sources generative systems pull from when they answer a question about your category. The mechanism is easy to break, and most companies have broken it without noticing.

A model does not watch your video

Start here, because it makes everything else make sense. When ChatGPT, Claude, Gemini or Perplexity surfaces a video as a source, no model has sat through it. There is no eye. What the system has is text: a transcript, a set of captions, a title, a description, chapter markers, and whatever other pages on the web have embedded that video or written about it.

That is the whole substrate. To a retrieval system, a video is a text document with a thumbnail attached, and everything a model says about it, it says because something turned the sound into words first. Which leads to a blunt conclusion: a video without a clean transcript is close to mute. You can have a beautifully produced explainer that converts on the landing page and contributes nothing at all to how a model describes your company, because the only text it emits is a title and a few lines of description written by someone in a hurry.

The transcript is the video, as far as a model is concerned

Automatic transcription is remarkable at ordinary language and unreliable at exactly the words you care about most. General vocabulary is easy: the system has heard "supply chain" countless times. Proper nouns are hard. Brand names, product names, the surname of your founder, the name of a niche standard in your industry. These are the tokens the transcriber invents, mangles, or splits in half.

And a brand name transcribed wrong attributes nothing to you. If the transcript renders your company as a common word or a made-up compound, then to any system reading that text your company was never mentioned. The video is talking about somebody else.

Check it yourself, right now

This is falsifiable in about the time it takes to make coffee.

  • Open one of your own videos on YouTube, the one where your brand is spoken most often.
  • Open the transcript panel: on desktop it sits under the description, on mobile behind the same menu.
  • Search that transcript for your brand name. Not with your eyes, with the browser's find function, typed exactly as you spell it.
  • Is it there, spelled correctly, every time it was said? Or does it appear as something else, or not at all?

Do the same for your product names and any technical term your category depends on. If the transcript is wrong, the fix is manual: upload a corrected caption file. YouTube accepts one and replaces the automatic guess with your text. That single upload is the difference between a video that mentions you and one that does not.

Title and description are not decoration

The title is the strongest short piece of text attached to the asset, and the description is the longest piece of prose you control sitting next to it. Both get read. Yet descriptions are still routinely written as a link dump: a sentence, then the social handles, then the timestamps. If the description does not state, in ordinary sentences, what the video answers and who made it, you have left the most retrievable field on the page empty and filled it with hyperlinks instead.

Chapters are structure, and retrieval likes structure

Chapter markers split a long video into named segments. To a human they are a convenience. To a retrieval system they are a set of labelled passages, each with its own slice of transcript underneath. A long webinar with no chapters is an undifferentiated wall of talk, and the part that answers a specific question is buried in the middle with nothing pointing at it. The same webinar with honest chapter titles becomes a small library of answerable units. Name them as the thing they cover, not as an inside joke: "Pricing and licensing" retrieves, "The fun part" does not.

Why a question in the title works so well

The format "ask the question in the title, answer it in the video" performs unusually well as a source, for a structural reason that has nothing to do with click-through rates. Retrieval works on similarity. A user asks a system something shaped like a question, the system goes looking for text that resembles that question, and the text that most resembles a question is a question. A title phrased the way a buyer would phrase it, followed by a transcript that answers it directly in the opening minute, is a close match for the query a real person types. A title phrased as a slogan matches nothing, because nobody has ever typed your slogan into anything.

It also matters where the answer sits. The passage that gets retrieved is the one where the answer is near the question, so a video that buries its payoff behind a long introduction buries the retrievable text too.

The pages around the video matter as much as the video

A video sitting alone on a platform is one document. A video embedded on your own site, next to a written summary, inside a page other people link to, is a cluster of documents pointing at the same thing. Models pick up video content through those surrounding pages constantly: a blog post that embeds the talk and paraphrases it, a documentation page linking to the demo, a forum thread quoting your founder. The video is the origin; the pages are the amplifier.

So the practical move is not to publish and walk away. Give every video that matters a written home: a page on your own domain with the embed, a real summary, and the transcript as readable text. That page can be retrieved even by systems that never touch the video platform.

What being a source does not get you

Being readable is a precondition, not a promise. A model can have your transcript sitting in its retrieved context and still write an answer naming three other companies. Nobody can guarantee that a video ends up cited, and anyone who tells you otherwise is describing a lever that does not exist. You control whether the machine has anything to work with. You do not control what it does with it.

What is actually worth measuring

The useful question is not how many views the video got. It is: when an AI answers a question about my category, do my videos appear among the sources it used?

Views cannot approximate that. Answering it means asking the buying questions across the systems people use, in each language you sell in, then reading the sources behind each answer instead of only the prose on top. Sometimes a video you had forgotten about turns out to be quietly doing the work. More often the sources are somebody else's pages, and none of your material is in the room.

That is the layer PSentry reports on: it runs your prompts across ChatGPT, Claude, Gemini and Perplexity, in each market, and shows which sources the models used to talk about your category, alongside where you were mentioned and who was named instead. It measures. It does not optimize your videos, and it will not promise you a citation.

Fixing a transcript is a small piece of work with a clear mechanism behind it. Whether it changes anything downstream is an empirical question, and the only way to know is to look at what the models cite before and after. Almost nobody looks.

Frequently Asked Questions

Do models really read YouTube transcripts?

They read text associated with the video, and the transcript is the largest piece of it. Systems that retrieve live pages can reach it directly; systems answering from training tend to have absorbed it through pages that quoted or embedded the video. Either way the route runs through words, not pixels.

Should I upload my own captions or trust the automatic ones?

Check the automatic ones first. If your brand and product names come through correctly, leave them. If they do not, upload a corrected file: a misspelled brand in a transcript attributes the content to nobody.

Does this apply to platforms other than YouTube?

The mechanism does: any video is legible to a model only through its surrounding text. What differs is how much of that text is publicly reachable. A video behind a login, or hosted somewhere that publishes no transcript, is invisible in retrieval terms however good it is.

If I fix everything, will my video get cited?

Maybe. Nobody can promise it, and a vendor who does is selling something they cannot deliver. Correct transcripts and clear titles make a video eligible to be used as a source. Whether a given answer uses it depends on the question, the model, the language, and what else exists on the web about your category.

How would I know if a video became a source?

By reading the sources attached to AI answers about your category, repeatedly, rather than by watching the view count. A single check tells you nothing: these systems are probabilistic, and two identical questions can produce two different sets of citations. Only a pattern, across many questions and over time, means anything.