Crawler policy
Responsible source retrieval
Operational limits used when primary-source pages are fetched for verification.
Identification
WorldSchoolIndexBot/1.0 identifies every automated request and links back to this policy.
Access controls
robots.txtis checked before any source page is requested.- Requests are limited to one every three seconds per domain, or a slower published crawl delay.
- No more than three school domains are contacted concurrently.
- Blocked, unreachable, and unsuccessful requests are recorded explicitly.
Purpose
Fetched pages are used to verify structured school fields and retain an auditable local source cache. Unsupported fields remain null.