Home · Data
Data, sources and our crawler
Every fact on this site carries its source, its license basis and the date it was last checked. This page lists where data comes from and what we never do.
Where the data comes from
| Tier | Source | License basis |
|---|---|---|
| 1 | Open data and public records | cdla-permissive-2.0, apache-2.0, odbl-1.0, public-record, official-calendar |
| 2 | The business's own website | robots-permitted-page |
| 3 | Licensed opinion source | licensed:<contract>, official-api:<name> |
| 4 | Cloute first-party | first-party-cloute |
| 5 | Stated by the business | business-supplied |
What we never do
- We do not scrape Google, Yelp, Tripadvisor, Booking, Expedia or any site whose terms forbid it.
- We do not bypass blocks, rate limits, logins or CAPTCHAs, and we do not disguise our crawler.
- We do not store or republish review text. We store counted aspects, sentiment and dates.
- We do not use photos from brand sites, travel sites or Google. Photos come from the business with rights.
- We do not let Google review text or any copied star rating feed a score.
- We do not sell inclusion, rank or score. Sponsored placements are labelled and never change a score.
Our crawler
User agent: Mozilla/5.0 (compatible; JourneyedBot/1.0; +https://journeyed.ai/data/crawler/)
- It identifies itself with the user agent above and obeys robots.txt, including Crawl-delay.
- It waits at least 10 seconds between requests to one site, fetches one page at a time, and reads at most 12 pages per site per refresh.
- It stops at a block, a login or a challenge page and does not retry with another agent, address or browser.
- To ask us to stop, block the crawler in robots.txt or write to hello@journeyed.ai.
Sources
- Overture Maps Places · cdla-permissive-2.0 · checked
Last checked