Dynamic bot rendering and bot routing: setup guide
Dynamic bot rendering, also called bot routing, sends crawlers a rendered HTML snapshot of each page while people get your JavaScript app as usual. This guide covers the routing rule itself: which user agents to route, what to leave alone, how to cache it safely, and how to test it. For the concept, see What is dynamic rendering.
If you use Encited, the setup guide for your host already includes a working rule. Use this page to understand what that rule does or to adjust it.
How the routing rule works#
Every request goes through one check before it reaches your site:
- Read the
User-Agentheader. - If it matches a crawler, return the rendered snapshot for that URL.
- Otherwise, pass the request to your host unchanged.
const CRAWLERS =
/googlebot|google-inspectiontool|bingbot|applebot|duckduckbot|gptbot|oai-searchbot|chatgpt-user|claudebot|claude-user|claude-searchbot|perplexitybot|perplexity-user|facebookexternalhit|twitterbot|linkedinbot|slackbot|discordbot|whatsapp|telegrambot/i;
export function isCrawler(request) {
const ua = request.headers.get("user-agent") ?? "";
return CRAWLERS.test(ua);
}
Which user agents to route#
| Group | User agents | Why |
|---|---|---|
| Search engines | Googlebot, Google-InspectionTool, Bingbot, Applebot, DuckDuckBot | Indexing. Include Google-InspectionTool so Search Console's URL Inspection sees what Googlebot sees. |
| AI crawlers | GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot, PerplexityBot, Perplexity-User | Most don't run JavaScript, so without routing they get an empty page. |
| Link previews | facebookexternalhit, Twitterbot, LinkedInBot, Slackbot, Discordbot, WhatsApp, TelegramBot | Titles, descriptions and images for shared links. |
Check your logs every few months for crawlers that aren't on the list. Any crawler missing from the rule gets your app shell.
Google-Extended doesn't belong here. It's a robots.txt token for Gemini training, and no crawler uses it as a user agent.
What not to route#
Only route page requests. Pass these through untouched:
- Static assets:
.js,.css, images, fonts, video. - API routes and anything under paths like
/api/. - Files that are already plain HTML or text:
robots.txt,sitemap.xml,/.well-known/*. - Logged-in areas, checkout and account pages.
A simple way to do this is to route only GET requests whose path has no file extension, plus any paths you list explicitly.
Verify that crawlers are real#
Anyone can send a request with a Googlebot user agent. If your snapshots contain nothing that differs from the public page, a fake crawler gets nothing special, so most sites don't need verification. If you want it anyway:
- Google and Bing: do a reverse DNS lookup on the IP, check that the hostname ends in
googlebot.com,google.comorsearch.msn.com, then do a forward lookup on that hostname and confirm it returns the same IP. - Published IP ranges: Google and OpenAI publish the IP ranges their crawlers use. Check each provider's crawler documentation for the current files.
Caching#
The same URL now has two versions, so caches need to keep them apart:
- Send
Vary: User-Agenton routed responses, or put a crawler flag in your CDN cache key. - Never cache the snapshot under the same key people use. If a CDN serves the snapshot to visitors, they get a page with no interactivity.
- Keep snapshot freshness separate from your CDN cache. Refresh snapshots when content changes, and on a schedule for sections that change often.
Status codes and redirects#
Return the status the page should have:
- A route that shows "not found" in the app should return
404to crawlers. - A redirect the app performs in the browser should be a
301or302for crawlers. - If the origin is down or returns an error, don't save that response as the snapshot.
Single-page apps return 200 for every URL by default, so this needs a signal from the page, such as a meta tag the app sets on its not-found route. With Encited, set <meta name="prerender-status-code" content="404"> on that route and crawlers get a 404.
Setup by host#
Encited setup guides include the rule above, adapted for each host:
- Cloudflare Workers
- Vercel
- Netlify Edge Functions
- AWS CloudFront with Lambda@Edge
- nginx
- Express
- No-code setup via DNS
For nginx, the matching part of the rule looks like this:
map $http_user_agent $is_crawler {
default 0;
"~*(googlebot|google-inspectiontool|bingbot|applebot|gptbot|oai-searchbot|chatgpt-user|claudebot|claude-user|perplexitybot|perplexity-user|facebookexternalhit|twitterbot|linkedinbot|slackbot)" 1;
}
Test the setup#
Compare responses. Fetch the same page as a crawler and as a browser:
curl -s -A "Googlebot" https://yoursite.com/page | wc -c curl -s -A "Mozilla/5.0" https://yoursite.com/page | wc -cThe crawler response should contain your page text and usually be larger than the browser response, which is the app shell.
Check status codes. Request a URL that doesn't exist with a crawler user agent and confirm it returns
404.Check Search Console. Use URL Inspection on a page and confirm the HTML Googlebot fetched includes your content.
Check assets. Confirm
.jsand.cssfiles return the same response for both user agents.Use the crawler simulator to see what each crawler receives.
Troubleshooting#
| Symptom | Likely cause |
|---|---|
| Visitors see a page that doesn't respond to clicks | The CDN is serving the snapshot to people. Check the cache key and Vary header. |
| Crawlers still get the empty shell | The user agent isn't in the rule, or the path is excluded. |
| An AI crawler gets the shell but Googlebot gets the snapshot | The rule only lists search engines. Add the AI user agents above. |
| Crawlers get an old version of the page | The snapshot hasn't been refreshed since the change. Re-render the page. |
| Crawlers get an error page | The snapshot was saved during an outage. Re-render and add a health check before saving. |
| Not-found pages return 200 to crawlers | The app has no not-found signal for the renderer to read. |
