Crawler policy

Responsible source retrieval

Operational limits used when primary-source pages are fetched for verification.

Identification

WorldSchoolIndexBot/1.0 identifies every automated request and links back to this policy.

Access controls

  1. robots.txt is checked before any source page is requested.
  2. Requests are limited to one every three seconds per domain, or a slower published crawl delay.
  3. No more than three school domains are contacted concurrently.
  4. Blocked, unreachable, and unsuccessful requests are recorded explicitly.

Purpose

Fetched pages are used to verify structured school fields and retain an auditable local source cache. Unsupported fields remain null.