Use SitemapKit to discover all URLs, then crawl them with Crawl4AI for LLM-friendly content extraction. The perfect combination for building AI training datasets.
Use the secret or environment-variable mechanism provided by your ai frameworks setup. Never expose the key in client-side code.
Adapt the request shown in the Crawl4AI example below and send the target domain from a trusted server process.
Handle discovered sitemap files, extracted URLs, lastmod values, truncation flags, and API errors before passing the result downstream.
import requests
from crawl4ai import AsyncWebCrawler
# Step 1: Get all URLs via SitemapKit
resp = requests.post(
"https://sitemapkit.com/api/v1/sitemap/full",
headers={"x-api-key": "YOUR_API_KEY", "Content-Type": "application/json"},
json={"url": "docs.example.com"}
)
urls = [u["loc"] for u in resp.json()["urls"]]
print(f"Found {len(urls)} URLs to crawl")
# Step 2: Crawl with Crawl4AI for LLM-ready content
async with AsyncWebCrawler() as crawler:
for url in urls[:50]:
result = await crawler.arun(url=url)
if result.success:
# result.markdown contains clean, LLM-ready text
save_to_dataset(url, result.markdown)Check authentication, endpoints, response fields, and limits.
Preview the URL extraction result before writing integration code.
See how teams use sitemap data in SEO and engineering workflows.
Look up lastmod, sitemap indexes, and other response concepts.
100 free API calls/month. No credit card required.
Use SitemapKit as a tool in LangChain agents or as a URL source for document loaders. Get all URLs from a domain to build RAG pipelines.
Feed sitemap URLs into LlamaIndex for building knowledge bases. Discover all pages on a domain, then index them for RAG-based question answering.
Build OpenAI agents that can discover and analyze website sitemaps. Give your agent the ability to map out any domain's content structure.