A search engine cannot see an image or watch a clip the way a person can. It reads the text attached to the file and the words around it, and infers the rest. That is the whole basis of image and file SEO: since the engine works from text, the text you attach to your media is what decides whether the media gets found at all. The last two articles covered the words on the page and the tags around your content, and this one goes down to the file itself, where a few small habits make your images and clips readable to search instead of invisible to it.

Two things ride on this. Image search is its own stream of traffic, with people finding a creator through a picture and following it back to the source, and well-described media also helps the page it sits on rank, because the file’s text reinforces what the page is about. Bare, unlabeled media does neither. It loads as a wall of pixels the engine cannot interpret, contributing nothing to the page and never surfacing on its own. The fix is to attach real text at every level the file offers.

Start with the filename, before the file is ever uploaded. A camera names a photo something like a string of letters and numbers that tells a search engine nothing, so rename it to a short descriptive phrase first, lowercase, with hyphens between the words, using the term someone would actually search. The filename is the first piece of text the engine reads about the image and a genuine ranking signal, and it follows the same descriptive logic as the clean URLs from the SEO fundamentals article. Rename in batches as part of your editing pass so it never becomes a separate chore.

Alt text is the heart of this article and the part worth getting right. The alt attribute is a written description of the image that lives in the page’s code, and it does two jobs at once. It is read aloud by a screen reader to a blind or low-vision visitor, an accessibility feature a real share of your audience depends on and one the accessibility chapter treats as a baseline. Search engines also use it to understand what the image shows. Write it as a true, natural description of what is actually in the frame, working the relevant keyword in only where it fits honestly. Alt text crammed with repeated keywords reads as spam to an engine and is useless noise to the person relying on a screen reader, so describe the image as you would to someone who cannot see it.

For this work, keep the alt description accurate and plain rather than turning it into erotic copy. Its job is to describe the image for accessibility and search, and gratuitously explicit alt text serves neither while giving the page’s content classifier more reason to filter it. A clear, matter-of-fact description does the work. Grouping your explicit images together in their own section of the site, as the SEO fundamentals article noted, also helps the search engines classify the page correctly instead of mislabeling your whole site.

The visible caption under an image and the text in the paragraphs around it tell the engine as much as the hidden attributes do, and they tell the reader more. A picture set inside real writing that explains or frames it carries more weight than the same picture dropped in alone, because the surrounding words confirm the topic. This is the same reason the fundamentals article pushed for text around media: the image and the words it sits among explain each other.

File size is the part of this that is easy to ignore and quietly costly. Large, uncompressed images make a page slow to load, slow pages rank worse, and visitors leave before a heavy page finishes drawing, which a content-heavy gallery feels hardest of all. Resize images to the dimensions they actually display at, compress them, and use the efficient modern image formats your site supports, since the standard formats shift over time and a current one is worth confirming. Lazy loading, where images load as a visitor scrolls to them rather than all at once, keeps a long page quick. A site SEO plugin handles much of this for you, so lean on it.

Video carries the same problem in a heavier form, since an engine can watch a clip even less than it can read a photo. What it can read is the title, the description, the captions or transcript, and the thumbnail, so those are where a video earns its visibility. Captions and a transcript give the engine actual crawlable text from inside the clip, on top of the accessibility they provide, which is why the subtitles work from the gear chapter and the captioning article pay off twice. The thumbnail is what earns the click once the clip surfaces, the same thumbnail craft the persona chapter covered. Marking your media up with the appropriate structured data helps search engines understand and feature it, and the specifics of that markup are worth checking against current guidance since they change.

Almost everything in this article serves accessibility and search in the same motion. Descriptive alt text, captions, and transcripts make your work usable by people who cannot see or hear it and legible to the engines at the same time, so the effort counts twice and the accessible version is also the discoverable one. The accessibility chapter goes deeper into making your whole site and profile work for everyone.

That closes out how people find you, through search, tags, and the files themselves. All of that discovery work is only worth doing if you can tell what is actually landing, which page ranks, which clip pulls traffic, which term converts. Seeing that clearly is the job of analytics, and the next part of the chapter starts there, with the handful of metrics actually worth watching.