Build OpenAI agents that can discover and analyze website sitemaps. Give your agent the ability to map out any domain's content structure.
Use the secret or environment-variable mechanism provided by your ai frameworks setup. Never expose the key in client-side code.
Adapt the request shown in the OpenAI Agents SDK example below and send the target domain from a trusted server process.
Handle discovered sitemap files, extracted URLs, lastmod values, truncation flags, and API errors before passing the result downstream.
from agents import Agent, Runner, function_tool
import requests
@function_tool
def discover_sitemaps(domain: str) -> dict:
"""Discover all XML sitemaps on a domain."""
resp = requests.post(
"https://sitemapkit.com/api/v1/sitemap/full",
headers={"x-api-key": "YOUR_API_KEY", "Content-Type": "application/json"},
json={"url": domain}
)
data = resp.json()
return {
"sitemaps_found": len(data["sitemaps"]),
"total_urls": len(data["urls"]),
"sample_urls": [u["loc"] for u in data["urls"][:10]],
}
agent = Agent(
name="SEO Analyst",
instructions="You analyze websites. Use discover_sitemaps to map out a domain.",
tools=[discover_sitemaps],
)
result = Runner.run_sync(agent, "Analyze the sitemap of vercel.com")
print(result.final_output)Check authentication, endpoints, response fields, and limits.
Preview the URL extraction result before writing integration code.
See how teams use sitemap data in SEO and engineering workflows.
Look up lastmod, sitemap indexes, and other response concepts.
100 free API calls/month. No credit card required.
Use SitemapKit as a tool in LangChain agents or as a URL source for document loaders. Get all URLs from a domain to build RAG pipelines.
Feed sitemap URLs into LlamaIndex for building knowledge bases. Discover all pages on a domain, then index them for RAG-based question answering.
Use SitemapKit to discover all URLs, then crawl them with Crawl4AI for LLM-friendly content extraction. The perfect combination for building AI training datasets.