Home ยท Help
Our crawler
Journeyed reads hotels' own websites to check facts such as check-in time, parking fees and pet rules. This page explains how our crawler behaves, what it reads and how to stop it.
User agent
Every request carries this user agent: Mozilla/5.0 (compatible; JourneyedBot/1.0; +https://journeyed.ai/data/crawler/). The product token that robots.txt groups match is JourneyedBot. We never rotate user agents or send requests under another name.
robots.txt
The crawler obeys robots.txt as RFC 9309 describes it. It follows the JourneyedBot group if there is one, otherwise the * group, and honors Crawl-delay. It reads robots.txt again at least every 24 hours.
A missing robots.txt (404 or 410) means the site allows crawling. If robots.txt fails with a server error or times out, the crawler fetches nothing from the site until it answers.
We are stricter than RFC 9309 on purpose: if robots.txt itself answers 401, 403 or 451, or shows a challenge page, we treat the whole site as closed for this refresh.
Pace
- At least 10 seconds between two requests to one site, or longer if Crawl-delay asks for it, and only 1 request at a time per site.
- At most 12 pages per site per refresh: the home page and the pages most likely to state facts, such as amenities, parking, policies, rooms and FAQ.
- A 15-second timeout and a 2 MB limit per page. Redirects are followed only within the same site, at most 3 in a row.
- On 429 or 503 the crawler waits as long as Retry-After asks, otherwise it backs off from 1 hour up to 7 days. After 5 failures in a row it leaves the site alone until the next refresh.
What it does not do
- It runs no JavaScript and uses no headless browser, no proxies and no cookies carried between sites.
- A block, a login or a challenge page ends crawling of that site for the refresh. It does not retry with another user agent, address or browser.
- It never crawls search engines, review sites, booking and travel sites, or social networks and forums.
What we extract
Facts only: things like check-in and check-out times, parking and pet fees, breakfast, shuttle, pool, accessible rooms and free Wi-Fi. Each fact is stored with the exact page it was read from and the date it was checked, and the site shows that source next to the fact.
We never copy guest reviews, ratings, photos or descriptive copy.
How to opt out
Add a group for the crawler to your robots.txt. The crawler stops fetching from the site the next time it reads robots.txt, within 24 hours:
User-agent: JourneyedBotDisallow: /
A Disallow rule for a single path keeps the crawler off just that part of the site. To have facts read from your site corrected or removed, write to hello@journeyed.ai.
Contact
For questions about the crawler, or to report a problem with it, write to hello@journeyed.ai.
Last checked