v1.0 — Now in public beta

Sitemap Extractor API
XML Parser & Scraper

Parse and scrape XML sitemaps, follow nested sitemap indexes, and extract clean URL lists with one REST API. The free API key includes 20 credits per month.

No credit card required · 20 credits/month included

Cached responses supported SSRF-safe by design Free tier included
API response
POST /api/v1/sitemap/full

{
  "url": "https://example.com"
}

→ 3 sitemaps discovered
→ 1,284 URLs extracted
→ deduplicated JSON response

Three endpoints. Full coverage.

Discover

Finds all sitemaps for a domain by parsing robots.txt and probing common sitemap paths.

POST /api/v1/sitemap/discover

Use the free XML Sitemap Finder

Prefer a manual check? Find and verify a sitemap manually.

Extract

Parses sitemap XML and recursively follows sitemap indexes up to 5 levels deep. Returns up to 50k URLs.

POST /api/v1/sitemap/extract

Count sitemap URLs and extract the list

Planning a scraper? Use sitemaps to build a crawl queue.

Full

Discover + Extract in one call. Finds all sitemaps, extracts every URL, and deduplicates the results.

POST /api/v1/sitemap/full

Read the API documentation

Sitemap scraper API or sitemap parser API?

The input decides which endpoint you need. Use /api/v1/sitemap/full when you only have a domain: it checks robots.txt and common sitemap paths before extracting the URLs. Use /api/v1/sitemap/extract when you already have a sitemap URL: it parses that file and follows any nested sitemap indexes.

Both return deduplicated URL records. See the request and response examples before choosing an endpoint.

Use the sitemap API in your stack

Extract sitemap URLs in GitHub Actions or a CLI

Run the same nested sitemap extraction from a terminal or save the JSON result as a workflow artifact. Both use the open-source SitemapKit CLI.

CLI
export SITEMAPKIT_API_KEY=sk_live_...
npx github:0nl1n1n/sitemapkit-cli full https://example.com
GitHub Actions
- name: Extract sitemap URLs
  id: sitemap
  uses: 0nl1n1n/sitemapkit-cli@v1
  with:
    api-key: ${{ secrets.SITEMAPKIT_API_KEY }}
    command: extract
    url: https://example.com/sitemap.xml

- uses: actions/upload-artifact@v4
  with:
    name: sitemap-urls
    path: ${{ steps.sitemap.outputs.result-file }}

Call the sitemap parser API from JavaScript or Python

Send a domain to the full endpoint when you want discovery and XML parsing in one request. Replace the sample key with your SitemapKit API key.

JavaScript
const response = await fetch("https://sitemapkit.com/api/v1/sitemap/full", {
  method: "POST",
  headers: {
    "x-api-key": process.env.SITEMAPKIT_API_KEY,
    "Content-Type": "application/json"
  },
  body: JSON.stringify({ url: "https://example.com" })
});

const result = await response.json();
console.log(result.data.urls);
Python
import os
import requests

response = requests.post("https://sitemapkit.com/api/v1/sitemap/full",
    headers={"x-api-key": os.environ["SITEMAPKIT_API_KEY"]},
    json={"url": "https://example.com"})
response.raise_for_status()

urls = response.json()["data"]["urls"]
print(urls)

Need the field definitions? Check the request and response contract.

Preview the API workflow

Discover sitemap files, follow nested indexes, and return a clean, deduplicated URL list in one request.

POST /api/v1/sitemap/full
{
  "url": "https://example.com"
}

→ discover sitemap files
→ follow nested sitemap indexes
→ extract and deduplicate URLs

One request. Every URL.

Send a domain, get back every URL from its sitemaps. No parsing XML yourself, no chasing sitemap indexes, no deduplication headaches.

  • Handles gzipped sitemaps, indexes, and nested indexes
  • Works with bot-protected and JavaScript-rendered sites
  • Response caching on paid plans (1h TTL)
terminal
$ curl -X POST https://sitemapkit.com/api/v1/sitemap/full \
  -H "x-api-key: sk_live_..." \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com"}'

{
  "success": true,
  "data": {
    "domain": "https://example.com",
    "sitemapsDiscovered": 3,
    "sitemapsProcessed": 3,
    "totalUrls": 1847,
    "truncated": false,
    "urls": [
      { "loc": "https://example.com/", "lastmod": "2026-02-20" },
      { "loc": "https://example.com/about", "lastmod": "2026-01-15" },
      ...
    ]
  }
}

Start extracting URLs via API

Free tier includes 20 credits/month. No credit card required.