WEBSITE CRAWLER

Connect your docs, keep them in sync

Add a URL. BeforeQuery discovers every page, extracts clean content, and automatically re-indexes when you publish updates. No code, no webhooks, no manual uploads.

How the crawler works

Built for the real-world complexity of documentation sites

Sitemap Discovery

BeforeQuery fetches sitemap.xml first, then follows links to discover every page the crawler should index. No manual URL lists required.

JavaScript Rendering

Uses a headless browser to render client-side content before extraction. SPAs, Next.js docs sites, and Docusaurus all index accurately.

Clean Markdown Extraction

Strips nav, footers, ads, and boilerplate. Preserves code blocks, tables, and heading hierarchy — the signal that matters for RAG retrieval.

Incremental Sync

Re-crawls on a schedule you control (hourly, daily, or on webhook trigger). Only re-indexes pages that changed, keeping costs and latency low.

INCREMENTAL SYNC

Your index stays current automatically

Stale documentation is worse than no documentation — users get wrong answers. BeforeQuery polls your site on your chosen schedule and re-indexes only the pages that changed, keeping every answer grounded in the latest content.

Sync frequency
Hourly, daily, or on-demand
Change detection
ETag + content hash
Webhook trigger
POST /projects/:id/sync
Crawl depth
Configurable, default 10 hops
CRAWL STATUS
Pages indexed2,847 / 3,100
2,801
Indexed
46
Pending
12
Updated
Last sync: 4 minutes ago

Works with every docs platform

If it renders in a browser and is publicly accessible, BeforeQuery can index it

Docusaurus
Mintlify
GitBook
VitePress
Nextra
ReadMe
Sphinx
MkDocs
Astro Starlight
Eleventy

Frequently Asked Questions

Common questions about the BeforeQuery website crawler

Yes. BeforeQuery respects robots.txt directives by default. You can also configure a custom crawl rate and explicit allow/deny URL patterns in your project settings.
Yes. Pro and Enterprise plans support authenticated crawling via session cookies or custom HTTP headers. For complex auth flows, use the API to upload content directly.
BeforeQuery renders JavaScript before extracting content. SPAs, Next.js sites, Docusaurus, and any other JS-rendered platform are fully supported.
Most documentation sites of 100–500 pages complete in under 2 minutes. Larger sites (1,000+ pages) typically finish within 10–15 minutes.
Yes. Configure URL patterns to exclude (e.g., /changelog/*, /blog/*) in project settings. You can also provide an explicit allow-list of URL prefixes to crawl.

Index your first docs site in minutes

Enter a URL, click crawl, and your knowledge base is ready.

Get Started Free