Connect your docs, keep them in sync
Add a URL. BeforeQuery discovers every page, extracts clean content, and automatically re-indexes when you publish updates. No code, no webhooks, no manual uploads.
How the crawler works
Built for the real-world complexity of documentation sites
Sitemap Discovery
BeforeQuery fetches sitemap.xml first, then follows links to discover every page the crawler should index. No manual URL lists required.
JavaScript Rendering
Uses a headless browser to render client-side content before extraction. SPAs, Next.js docs sites, and Docusaurus all index accurately.
Clean Markdown Extraction
Strips nav, footers, ads, and boilerplate. Preserves code blocks, tables, and heading hierarchy — the signal that matters for RAG retrieval.
Incremental Sync
Re-crawls on a schedule you control (hourly, daily, or on webhook trigger). Only re-indexes pages that changed, keeping costs and latency low.
Your index stays current automatically
Stale documentation is worse than no documentation — users get wrong answers. BeforeQuery polls your site on your chosen schedule and re-indexes only the pages that changed, keeping every answer grounded in the latest content.
Works with every docs platform
If it renders in a browser and is publicly accessible, BeforeQuery can index it
Frequently Asked Questions
Common questions about the BeforeQuery website crawler
Index your first docs site in minutes
Enter a URL, click crawl, and your knowledge base is ready.
Get Started Free