Feed sitemap URLs into LlamaIndex for building knowledge bases. Discover all pages on a domain, then index them for RAG-based question answering.
Use the secret or environment-variable mechanism provided by your ai frameworks setup. Never expose the key in client-side code.
Adapt the request shown in the LlamaIndex example below and send the target domain from a trusted server process.
Handle discovered sitemap files, extracted URLs, lastmod values, truncation flags, and API errors before passing the result downstream.
import requests
from llama_index.core import VectorStoreIndex
from llama_index.readers.web import SimpleWebPageReader
# Get all URLs from sitemap
resp = requests.post(
"https://sitemapkit.com/api/v1/sitemap/full",
headers={"x-api-key": "YOUR_API_KEY", "Content-Type": "application/json"},
json={"url": "docs.example.com"}
)
urls = [u["loc"] for u in resp.json()["urls"]]
# Load and index pages
documents = SimpleWebPageReader(html_to_text=True).load_data(urls[:100])
index = VectorStoreIndex.from_documents(documents)
# Query the index
query_engine = index.as_query_engine()
response = query_engine.query("How do I configure authentication?")
print(response)Check authentication, endpoints, response fields, and limits.
Preview the URL extraction result before writing integration code.
See how teams use sitemap data in SEO and engineering workflows.
Look up lastmod, sitemap indexes, and other response concepts.
100 free API calls/month. No credit card required.
Use SitemapKit as a tool in LangChain agents or as a URL source for document loaders. Get all URLs from a domain to build RAG pipelines.
Use SitemapKit to discover all URLs, then crawl them with Crawl4AI for LLM-friendly content extraction. The perfect combination for building AI training datasets.
Build OpenAI agents that can discover and analyze website sitemaps. Give your agent the ability to map out any domain's content structure.