Apache Nutch logo

Apache Nutch

Extensible, scalable Java web crawler for large-scale data acquisition

Repository activity
  • Stars3.3k
  • Forks1.3k
  • Open Issues12
apache-nutch health score - Linux Foundation Insights
License

Apache-2.0

Languages
  • Java
  • HTML
  • Shell
Apache Nutch screenshot

About Apache Nutch

Apache Nutch is a mature, production-grade web crawler for collecting data from the web at scale. Written in Java, it runs crawl jobs on Apache Hadoop and integrates with search indexes such as Apache Solr.

Crawling is organized into configurable workflows rather than a fixed model, and a plugin architecture lets you adapt fetching, parsing, indexing, and scoring to different acquisition needs. This makes it suitable for everything from focused crawls to broad, distributed collection.

Nutch is an Apache Software Foundation project, released under the Apache License 2.0. It is self-hosted, with no hosted service, and is governed through Apache's community process.

Key features

  • Distributed crawling on Apache Hadoop
  • Configurable crawl workflows
  • Plugin architecture for fetch, parse, and index
  • Indexing into Apache Solr and others
  • Fine grained, per-deployment configuration

Details

First released
2009
Deployment
Self-hosted
Language
Java
Stack
Apache Hadoop · Solr
Governance
Apache Software Foundation
License
Apache-2.0