Stage 1 of 9
Crawl - where addresses come from
Nothing is crawled just because it exists. A host enters the queue only as an owner-approved seed, a reviewed public submission, a link found on a page we were already allowed to read, or a sitemap we have validated. Broad, uncontrolled crawling of the open web is switched off.
Discovery normalises every URL before deduplication, and bounds depth, query-parameter combinations, path repetition, calendar ranges and session-like addresses so that one site cannot generate unlimited work.
- URL state
discovered- Module
frontier