Robots.txt: the file you never think about until a crawler eats your budget
robots.txt is twenty lines at most, sits at the root of every domain, and most developers have never opened the one on a project they deployed. It has almost no logic in it: user-agent declarations, allow/disallow rules, and a pointer to your sitemap. Still, it decides two things: whether pages get crawled, and whether the crawler spends its time on your products or on your calendar pages. Wrong in either direction, and your SEO quietly gets worse with no error message anywhere.
The five directives worth knowing
User-agentstarts a block.*means all crawlers; specific names likeGooglebotoverride it.Disallowblocks a path. An emptyDisallow:means allow everything, which is a common way to write "I don't care" but reads as a mistake later.Allowoverrides a Disallow for a specific path, useful when you block a whole directory but want one public page inside it.Sitemap:is where the crawler finds your sitemap. Absolute URL, on the same domain. This one is invisible in most setups and is the line people forget.- A blank line starts a new
User-agentblock. One unsplit block is the second mistake people make: two rules separated by a blank line are two separate rule sets, not a refinement of the first.
Four real mistakes
Blocking what a page needs to render. Disallow /assets/ and Googlebot might still fetch your HTML, but it can't get your CSS or images, so it renders your page looking broken and your rich results get capped. This is the single most common robots.txt mistake I see on real projects. If you have a /build/ or /dist/ directory next to a disallowed admin path, leave those alone.
Blocking admin paths that were already public. If you deployed a staging environment without auth and then robots-disallowed it, the URL is already known and has already been linked. Disallow stops crawling, not indexing. What you actually want is noindex on the page, or a login.
Blocking the page you want indexed. This one is ironic. A dev blocks /checkout/ without checking that /products/checkout/ matches that pattern first and then watches organic revenue go down. Read every row. The patterns are prefix-based, not path-segment-based, and a leading slash does not make a pattern narrower than you'd expect.
Forgetting that robots.txt is per hostname. https://www.yourdomain.com/robots.txt does not cover https://yourdomain.com/robots.txt. If both hostnames resolve, the crawler reads two different files and indexes from either one.
Validate it before you deploy
Google Search Console has a robots.txt tester (in the legacy tool it was called just that, and it shows the parsed result of any URL). If you're not in GSC for this project, a robots.txt generator with a preview will show you what a crawler reads. The key thing it surfaces is pattern matching: does Disallow: /admin also catch /administrator or /admin-panel? Yes, in most implementations, because it's a prefix match. If you want to exclude /administrator, you add that as an explicit second rule, not as a clever regex.
One more thing worth checking: your staging environment. If you deploy to staging.yourdomain.com with the same robots.txt as prod, every crawler you're testing with will index the staging version. The fix is a single User-agent: * block with Disallow: / at the top of the staging file. Some teams skip it because "the file's small". It's small because nobody has looked at it.