Extensible, scalable Java web crawler for large-scale data acquisition
License
Apache-2.0
Languages
- Java
- HTML
- Shell

About Apache Nutch
Apache Nutch is a mature, production-grade web crawler for collecting data from the web at scale. Written in Java, it runs crawl jobs on Apache Hadoop and integrates with search indexes such as Apache Solr.
Crawling is organized into configurable workflows rather than a fixed model, and a plugin architecture lets you adapt fetching, parsing, indexing, and scoring to different acquisition needs. This makes it suitable for everything from focused crawls to broad, distributed collection.
Nutch is an Apache Software Foundation project, released under the Apache License 2.0. It is self-hosted, with no hosted service, and is governed through Apache's community process.
Key features
- Distributed crawling on Apache Hadoop
- Configurable crawl workflows
- Plugin architecture for fetch, parse, and index
- Indexing into Apache Solr and others
- Fine grained, per-deployment configuration
Details
- First released
- 2009
- Deployment
- Self-hosted
- Language
- Java
- Stack
- Apache Hadoop · Solr
- Governance
- Apache Software Foundation
- License
- Apache-2.0
