SEO
XML Sitemap and robots.txt: A Plain-English Setup Guide
By Kavin P · · 7 min read

Two small files quietly shape how search engines explore your website. One is a map that lists the pages you want found. The other is a set of polite instructions about where crawlers should not go. Used well, the XML sitemap and robots.txt help search engines spend their time on your best pages. Used carelessly, they can hide the pages you care about most.
The Roles in One Minute
Picture a visitor arriving at a large building.
- The XML sitemap is the directory in the lobby. It lists the rooms you want visitors to find.
- The robots.txt file is the sign on certain doors saying "staff only, please do not enter".
The directory suggests; the sign requests. Neither guarantees how a search engine behaves, but well-behaved crawlers pay attention to both. Understanding the difference between suggesting and requesting prevents most beginner mistakes.
XML Sitemaps: Your Page Directory
An XML sitemap is a file, usually found at an address like your domain followed by a sitemap file name, that lists URLs you would like search engines to know about. It can also include optional details such as when a page was last modified.
What Belongs in a Sitemap
Include pages that:
- Return a normal, successful response
- Are the preferred (canonical) version of the content
- You actually want people to find in search
- Offer unique value
Leave out redirected pages, error pages, duplicate versions, internal search results, thank-you pages, login areas and thin tag archives. A sitemap full of rubbish teaches search engines to trust it less.
This is where canonical tags matter. The address in your sitemap should match the canonical address of the page. If they disagree, you are sending mixed signals.
Why a Sitemap Helps
Search engines discover pages mainly by following links. A sitemap is a second route, useful when:
- Your site is new and has few incoming links
- You have many pages and some are deeply buried
- Pages are poorly linked internally
- Your content changes often
A sitemap does not force indexing and does not improve rankings by itself. It helps discovery. If a page is missing from search results, the sitemap is only one of several things to check.
How to Create One
- Use your platform. Most content management systems and SEO plugins generate a sitemap automatically. WordPress sites commonly have this built in or through a plugin. See SEO for WordPress websites.
- Check what it includes. Open the file in your browser and scan the list. Does it contain pages you would rather hide?
- Split large sites. Very large sites use an index file that points to several smaller sitemaps, for example one for posts, one for products and one for pages.
- Submit it. Add the sitemap address in your search performance tool so you can see how many of its pages were processed. Google Search Console basics explains where.
- Keep it current. Pages added or removed should appear or disappear automatically. Verify this after changes.
robots.txt: Your Crawling Instructions
The robots.txt file lives at the root of your domain and gives crawlers instructions about which areas they may or may not request. It is plain text, with simple rules.
A typical file contains:
- A line naming which crawler the rules apply to (an asterisk means all)
- Lines that disallow specific folders or patterns
- Optionally, lines that explicitly allow exceptions
- A line pointing to your sitemap address
You can mention your sitemap inside robots.txt, which is a tidy way to let any crawler find it.
What to Block (and What Not To)
Reasonable things to keep crawlers away from:
- Admin and login areas
- Internal search result pages that create endless combinations
- Cart and checkout steps
- Temporary staging folders (though a password is safer)
Be careful with these:
- Do not block CSS and script files that pages need to display properly. Search engines render pages, and blocking resources can make them misjudge your layout.
- Do not block pages you want indexed. This sounds obvious, but a development setting carried over to a live site is a classic cause of a vanished website.
- Do not rely on robots.txt for privacy. The file is public, and a blocked address can still appear in results if other sites link to it. Use passwords or proper access controls for anything sensitive.
Blocking Versus Removing From Search
This is the most misunderstood point. Blocking crawling is not the same as keeping a page out of search results.
- If a page is blocked in robots.txt, crawlers cannot read it, so they cannot see a "noindex" instruction on that page.
- To keep a page out of results, allow crawling and add a noindex instruction on the page, or use proper access control.
Choose the method that matches your goal, and do not stack conflicting methods on the same page.
How the Two Files Work Together
The files should tell a consistent story.
- Pages in your sitemap should not be blocked by robots.txt.
- Pages blocked by robots.txt should not appear in your sitemap.
- Pages with a noindex instruction should not be in the sitemap.
- Sitemap addresses should be canonical, live and indexable.
When these four rules hold, search engines get a clear, trustworthy picture of your site. Many problems flagged in reports come from breaking one of them.
A Hypothetical Walkthrough
Imagine a small online furniture shop. Its site has product pages, category pages, a blog, a cart, an account area and thousands of filter combinations.
A tidy setup might be:
- Sitemap lists the product pages, main categories, blog posts and key information pages.
- Robots.txt disallows the cart, account area and internal search results.
- Filtered category addresses are handled with canonicals or other controls, as covered in faceted navigation SEO.
- The sitemap address is listed in robots.txt and submitted in the search performance tool.
After launch, the owner checks the sitemap report monthly for pages discovered but not indexed, and investigates a handful at a time.
Common Mistakes
- Leaving a "discourage search engines" setting switched on after launch. Always test this before and after a go-live.
- Disallowing the entire site. A single misplaced slash can block everything. Review the file carefully whenever you edit it.
- Listing redirecting or broken addresses in the sitemap. Clean them up after every migration or redesign, using what you learned in 301 vs 302 redirects.
- Forgetting the file entirely. A missing sitemap is not a disaster, but a missing robots.txt can cause odd behaviour on some servers. Make sure a valid one exists.
- Using different versions of the domain. The sitemap should use the same preferred address as your canonicals and redirects.
- Treating the sitemap as a ranking tool. It is a map, not a boost.
Testing and Monitoring
- Open both files in a browser and read them like a visitor would.
- Use the testing and reporting features in your search performance tool to see whether specific addresses are blocked or excluded.
- Look at the sitemap report for errors, and at the indexing report for pages marked as excluded, then judge whether each exclusion is intentional.
- After any big change, such as a redesign or platform move, repeat these checks. The website redesign checklist is a good place to add them.
- Include these files in your regular technical SEO audit.
Special Cases Worth Knowing
Multilingual and International Sites
If you serve several languages or regions, the sitemap can help organise them, and each language version should be reachable and canonical in its own right. The principles in international SEO basics apply.
Images and Videos
Some sites add dedicated sitemaps for images or videos to help search engines discover media that is loaded in complex ways. These are optional and usually only worthwhile for media-heavy sites.
Very Small Sites
If your site has only a handful of well-linked pages, a sitemap is still easy to generate and does no harm, but it is not urgent. Make sure basic navigation links every page first.
Takeaway: Keep the Map Clean and the Doors Clear
The sitemap lists what you want found. The robots file explains where crawlers should not wander. Keep the first limited to good, canonical, indexable pages, use the second sparingly and carefully, and check both after every major change.
If you would like a hand reviewing these files on your own website, explore the free resources or contact Kavin.
Frequently asked questions
Does every website need an XML sitemap?
Not strictly, but it is helpful for most sites, especially new ones or those with many pages. A sitemap aids discovery, while good internal linking remains the main way search engines find your content.
Does robots.txt keep a page out of search results?
Not reliably. It asks crawlers not to visit a page, but the address can still appear if other sites link to it. To keep a page out of results, use a noindex instruction or access controls.
Where do I put my sitemap address?
List it in your robots.txt file and submit it in your search performance tool. Both steps are quick and help crawlers find the sitemap and help you monitor how its pages are processed.
How often should I check these files?
Review them after any redesign, migration, plugin change or new section launch, and include them in a regular technical check every few months. A small error in either file can have a large effect.
Related articles

SEO8 min read
301 vs 302 Redirects: Which to Use and When
301 vs 302 redirects explained with clear examples: when a move is permanent, when it is temporary, and how to avoid chains, loops and lost visitors.

SEO7 min read
Backlink Audit: How to Review Your Links Step by Step
A backlink audit shows who links to you, which links help and which need action. Follow this step-by-step method to review your link profile with confidence.

SEO7 min read
Canonical Tags Explained: Fix Duplicate Pages Calmly
Canonical tags explained in plain English: what they do, when to use them, common mistakes and a simple way to check that your preferred URLs are signalled.
