Internet Archive's web-scale, archival-quality web crawler
License
Other
Languages
- Java
- HTML
- Rich Text Format

About Heritrix
Heritrix is the Internet Archive's open-source, archival-quality web crawler. Written in Java, it captures web content at scale into WARC files for long-term preservation by researchers and archives.
It respects robots.txt and META nofollow directives, and operators can tune politeness policies, identify crawls via the User-Agent, and configure crawl jobs in fine detail. A web admin interface and a REST API drive crawl operation and monitoring.
Heritrix is maintained by the Internet Archive and released under the Apache License 2.0, with some bundled components under other licenses. It is self-hosted and also distributed as a Docker image.
Key features
- Web-scale crawling into WARC archive files
- Respects robots.txt and META nofollow tags
- Configurable politeness policies per crawl
- Web admin interface and REST API
- Detailed crawl job configuration
Details
- First released
- 2011
- Language
- Java
- Output
- WARC archive files
- Deployment
- Self-hosted · Docker
- Governance
- Internet Archive
- License
- Apache-2.0
