Heritrix logo

Heritrix

Internet Archive's web-scale, archival-quality web crawler

Repository activity
  • Stars3.2k
  • Forks788
  • Open Issues33
internetarchive-heritrix3 health score - Linux Foundation Insights
License

Other

Languages
  • Java
  • HTML
  • Rich Text Format
Heritrix screenshot

About Heritrix

Heritrix is the Internet Archive's open-source, archival-quality web crawler. Written in Java, it captures web content at scale into WARC files for long-term preservation by researchers and archives.

It respects robots.txt and META nofollow directives, and operators can tune politeness policies, identify crawls via the User-Agent, and configure crawl jobs in fine detail. A web admin interface and a REST API drive crawl operation and monitoring.

Heritrix is maintained by the Internet Archive and released under the Apache License 2.0, with some bundled components under other licenses. It is self-hosted and also distributed as a Docker image.

Key features

  • Web-scale crawling into WARC archive files
  • Respects robots.txt and META nofollow tags
  • Configurable politeness policies per crawl
  • Web admin interface and REST API
  • Detailed crawl job configuration

Details

First released
2011
Language
Java
Output
WARC archive files
Deployment
Self-hosted · Docker
Governance
Internet Archive
License
Apache-2.0