Home › Blog › Getting found on Google

Your Sitemap and Robots File, in Plain Words

A sitemap and a robots file are small parts of your site, but they can affect what Google finds and what it skips. If you run a local service business, this guide shows you what each file does, how to check yours, and the mistakes that can quietly block pages from search.

A small office desk where a website manager and a contractor review printed website notes beside a laptop, with a coffee mug, a notepad, and a phone on the table.

What these two files actually do

A sitemap is a list of pages you want search engines to know about. Think of it as a clean map of your site: home page, service pages, location pages, blog posts, and anything else you want Google to find faster and more reliably.

A robots file is a set of instructions for search crawlers. It can tell them where they may look and where they should stay out. Used well, it helps search engines focus on the pages that matter. Used badly, it can block the very pages you want customers to see.

  • A sitemap helps Google discover pages.
  • A robots file helps guide crawling.
  • Neither file makes a page rank by itself.

How to find yours without hunting around

Most sites keep the sitemap at yourdomain.com/sitemap.xml. Some sites use more than one sitemap, with a main file that points to other files for pages, posts, or images. If you do not know the address, open your site in a browser and try that path first.

The robots file is usually at yourdomain.com/robots.txt. Open it in a browser and read it like plain text. If you use WordPress, many sites create both files through the platform or an SEO plugin, but the exact setup can vary.

If the file opens in the browser, you can still inspect it without any special tools. You are looking for simple signs: does the sitemap location appear, and does the robots file avoid blocking important folders or pages?

  • Try /sitemap.xml and /robots.txt first.
  • If you use WordPress, check both files after site changes.
  • Open the files in a browser before changing anything.

What a healthy sitemap looks like

A healthy sitemap lists live pages that you want indexed. That means active service pages, city pages that are real and useful, contact pages, and other pages that add value to a customer. It should not be stuffed with broken links, duplicate pages, tag pages, test pages, or old drafts.

If a page shows in the sitemap, it should also be something you want search engines to visit. That sounds simple, but many sites include pages that should not be there. When that happens, you send mixed signals and make it harder for Google to read the site cleanly.

A good test is this: if a customer landed on that page, would it help them pick up the phone, book a visit, or understand your service? If the answer is no, the page may not belong in the sitemap.

  • Keep only useful live pages in the sitemap.
  • Remove test pages, tag pages, and old duplicates.
  • Check that the listed URL matches the live page exactly.

What a healthy robots file looks like

A healthy robots file is usually simple. It should not block your main site pages, your service pages, or your location pages. It may point search engines to your sitemap and leave the rest open unless there is a real reason to keep out private folders.

The common mistake is blocking too much. A single bad rule can hide sections of your site from Google. Another mistake is using the robots file to hide pages that should really be handled in another way, such as with a noindex tag or a redirect. The robots file is not the place for guessing.

If you see lines that block the whole site, a whole folder, or page paths that include your main services, pause before you save any changes. Even a small edit can have a big effect.

  • Do not block service pages by accident.
  • Use the robots file for crawling rules, not cleanup.
  • If a folder is private, keep it private for a reason.

The mistakes that quietly hide a site

One common problem is a sitemap that lists pages you deleted. If those links still sit in the file, Google may keep wasting time on dead pages. Another problem is a robots file that blocks pages before they can be crawled, which can keep them out of search even when the pages are live and useful.

A third issue is mixing old and new site structures after a redesign. You may move pages, change slugs, or switch themes, then forget to update the sitemap and robots file. That can leave broken paths in the sitemap or rules that no longer fit the site.

You can also hide pages by mistake when you copy a staging setup into the live site. Staging files often use blocking rules on purpose. If those rules follow the site live, Google may be shut out without anyone noticing right away.

  • Remove deleted pages from the sitemap.
  • Check for old block rules after a redesign.
  • Never move staging rules to the live site without review.

How to review both files this week

Start with the sitemap. Open it and scan for pages that should not be there. Look for broken links, old service pages, duplicate pages, tag archives, and test content. If the sitemap is very long, focus on the pages that bring customers to your business first.

Then open the robots file. Look for any rule that blocks the whole site or a key folder. Check that the sitemap location is listed near the top or bottom, where it can be found easily. If you use Google Search Console, you can also use it to spot crawling and indexing problems after you make changes.

If you are not sure what a rule does, copy it into a note and ask the person who built the site before you change it. A few minutes of caution can save you from hiding pages you worked hard to build.

  • Read the sitemap like a list of pages you are vouching for.
  • Read the robots file like a set of access rules.
  • Use Google Search Console to check for crawling issues after edits.

Questions owners ask

Do I need both a sitemap and a robots file?

Most business sites should have both. The sitemap helps search engines find the pages you want seen, and the robots file helps guide where crawlers can go. You can have one without the other, but together they give you more control.

Should I add every page on my site to the sitemap?

No. Add pages you want search engines to crawl and index, and leave out pages that are private, thin, duplicate, or temporary. If a page does not help a customer, it may not belong there.

Can the robots file keep a page out of Google forever?

It can keep search engines from crawling a page, which can stop it from being found or fully understood. But if a page is already discovered in other ways, blocking it is not always the right fix. For pages you want removed from search, the better choice is often a different setup, not a blanket block.

Want us to do it for you?

One message is enough

Tell us your trade and the towns you actually drive to. A person reads it and a person replies, usually the same day, in English or Spanish.

(561) 595-8715