How to find a competitor's new pages from their sitemap
Step by step: find a competitor's sitemap, save a dated copy, compare two copies and spot new comparison, alternative and industry pages.
By DanPublished 6 min read
A competitor's new landing pages say a lot about where they are heading. A page called "alternative to" your product means they are going after your customers. A page for a new industry means a new market. A new integrations page means a new partner. These pages often go live quietly, without a blog post, and sometimes without a link from the main menu.
Most of them still end up in one place: the site's sitemap. This guide shows how to find a competitor's sitemap, keep dated copies, compare them and pick out the pages worth reading. Once it is set up, it takes a few minutes a week.
What a sitemap is
A sitemap is a file a website publishes so search engines can find its pages. It is usually XML and lists one address per page, often with the date the page last changed. It looks like this:
<urlset>
<url>
<loc>https://example.com/pricing</loc>
<lastmod>2026-09-12</lastmod>
</url>
<url>
<loc>https://example.com/vs/rival</loc>
<lastmod>2026-09-29</lastmod>
</url>
</urlset>
Sitemaps are public on purpose. Anyone can open one in a browser, and reading it is the same as reading any other public page.
Step 1: find the sitemap
Try these, in order, on the competitor's domain:
/robots.txt. Openhttps://their-domain.com/robots.txtand look for lines that start withSitemap:. This is the most reliable way, because the site names its own sitemap there./sitemap.xml. The most common address./sitemap_index.xmlor/sitemap-index.xml. Common on sites with several sitemaps.
If none of these work, search for site:their-domain.com sitemap in a search engine. Also check subdomains like blog. or docs.: they often have their own sitemap.
Some sitemaps end in .xml.gz. That is the same file, compressed. Your browser may download it instead of showing it; unzip it and open the result.
Step 2: open every part of a sitemap index
Larger sites split their sitemap into several files and list them in a sitemap index. You can tell one apart because it contains <sitemap> entries instead of <url> entries. Each entry points to another sitemap, for example one for pages, one for blog posts, one per language.
Open each child sitemap and note which ones matter. For finding new landing pages, the sitemap for pages matters most. Blog post sitemaps are worth a look too, but they change often and are easy to follow in other ways, like an RSS feed.
Step 3: save a dated copy of the address list
You only see what is new if you have something to compare with. Save the list of addresses once a week, on the same day, with the date in the file name.
Without code. Open the sitemap in your browser, select everything, copy it and paste it into a new column in a spreadsheet, with the date as the column heading. Filter the column to rows that start with https so only the addresses are left. Repeat for each child sitemap that matters.
With a terminal (macOS or Linux). This saves a sorted list of every address in one sitemap:
curl -s https://example.com/sitemap.xml \
| grep -o '<loc>[^<]*' \
| sed 's/<loc>//' \
| sort -u > pages-2026-10-01.txt
For a .xml.gz sitemap, add | gunzip straight after the curl part. Run it once per child sitemap that matters, with a different file name for each.
Keep the request count low. One download per sitemap per week is enough for this job.
Step 4: compare two copies
In a spreadsheet, put last week's list in column A and this week's in column B. In column C, next to each address in B, add:
=IF(COUNTIF(A:A, B2)=0, "new", "")
Fill it down and filter column C to "new". Swap the columns in the formula to see pages that were removed.
In a terminal, comm compares two sorted files:
# New: in this week's list only
comm -13 pages-2026-09-24.txt \
pages-2026-10-01.txt
# Removed: in last week's list only
comm -23 pages-2026-09-24.txt \
pages-2026-10-01.txt
The first time, you have nothing to compare with. If the sitemap has <lastmod> dates, sort by them, newest first, and read the top of the list. That shows what changed recently, as long as the site fills in the dates honestly.
Step 5: spot the pages that matter
Most new addresses are blog posts, help articles and job ads. Skim past them and look for these patterns in the address:
- Comparison pages:
/vs/,-vs-,/compare/,/versus/. A page with your name in it means they are selling against you directly. - Alternative pages:
/alternatives/,-alternative,alternative-to-. Same as above, often aimed at people searching for a way out of a product. - Industry and use-case pages:
/industries/,/solutions/,/use-cases/,/for/. A new one points at a market they want to win. - Integration pages:
/integrations/,/apps/,/marketplace/. A new partner can change who they sell to. - Customer stories:
/customers/,/case-studies/. Shows which kinds of companies they are winning. - Campaign pages:
/lp/,/go/,/offer/,/webinar/. Often tied to a launch or a paid campaign. - Pricing and plans: a new page under
/pricing/, or a new plan name in an address. - New languages: a new
/de/,/fr/or/es/section usually means a push into a new country.
Removed pages matter too. A deleted plan page or a removed industry page can mean they stepped back from something.
Step 6: read the page and write it down
An address only tells you a page exists. Open each one that looks important and read it. Then write one line in a shared log, with the same fields every time:
Date found: 2026-10-01
Competitor: Northwind CRM
Page: /vs/your-product
Type: comparison
What it says: claims faster setup and
a cheaper entry plan than ours
What we do: check the claims, update
the battlecard, brief sales on Monday
Northwind CRM is a made-up company. If you keep battlecards, a new comparison page is the clearest sign that one needs an update. Our battlecard template has a section for exactly that.
What the sitemap won't show you
A sitemap is a useful signal, not a complete list. Keep these limits in mind:
- Not every page is listed. Pages built for ads are often left out on purpose and marked so search engines skip them. Some sites list only blog posts, or nothing at all.
- Some sitemaps are out of date. If it is generated by hand or by a forgotten plugin, new pages may never appear.
<lastmod>is often wrong. Some sites set every page to today's date on every build, which makes the field useless.- A listed page isn't always live. Test pages, drafts and pages that redirect sometimes slip in. Open the page before you act on it.
- Big sites are noisy. A site with thousands of help articles makes the weekly list long. Filter to the child sitemaps and patterns above.
When a page matters and the sitemap missed it, you will usually still find it in the site's main menu, footer or changelog, so check those once a month as well.
Let a tool do the weekly part
The manual routine works for two or three competitors. Past that, the weekly download and compare step is what gets skipped.
Plainrival's website change tracking reads your competitors' sitemaps every day, so a new page shows up within a day, alongside changes to the pages you already watch. The few new pages and changes that matter go into the weekly brief each Monday, with why each one matters and one suggested move.
Start free and add the competitors whose sitemaps you would otherwise check by hand.