# Apache Nutch

**Canonical:** https://apis.io/providers/apache-nutch/  
**APIs profiled:** 7

Apache Nutch is a highly extensible and scalable open-source web crawler software project built on Apache Hadoop data structures for batch processing. It provides a pluggable architecture supporting custom parse filters, scoring filters, index writers, and protocol implementations. Nutch integrates with Apache Solr and Elasticsearch for full-text search and exposes a REST API for managing crawl jobs, configurations, seed lists, and database queries. Governed by the Apache Software Foundation under the Apache License 2.0.

## Kin Score — 44.0 / 100 (developing)

Scored 2026-08-20 under rubric 0.12.0. Trend: flat (+0.0 from 44.0).

| Facet | Score |
|---|---|
| Discoverability | 64.8 |
| Contract Quality | 68.9 |
| Governance | 9.8 |
| Contract Governance | 9.8 |
| Operational Transparency | 36.8 |
| Developer Ergonomics | 45.2 |
| Commercial Clarity | 26.3 |
| Access Clarity | 26.3 |

## Agent readiness — 32.1 (agent-aware)

| Dimension | Value |
|---|---|
| Spec Presence | yes |
| Agentic Access | derived |
| Reversibility Documented | no |
| MCP Server | no |
| Auth Clarity | yes |
| Idempotency | no |
| Error Semantics | no |
| OpenAPI Examples | partial |
| Rate Limit Signal | documented |
| Event Surface Described | no |
| Agent Skills | no |
| Well Known Catalog | no |
| Consent Identity | no |
| Agent Card | no |
| Dry Run Mode | no |

## Access

Freemium · Self-serve signup — onboarding: self-serve, pricing: freemium, trial: no (confidence: high).

## APIs (7)

- **Apache Nutch Admin API** — Server administration operations
- **Apache Nutch Configuration API** — Manage Nutch configurations
- **Apache Nutch Database API** — Query the CrawlDB and FetchDB
- **Apache Nutch Job API** — Manage crawl jobs
- **Apache Nutch Reader API** — Read sequence files and webgraph data
- **Apache Nutch Seed API** — Manage seed URL lists
- **Apache Nutch Services API** — Auxiliary service operations such as CommonCrawl data dumps

## Agentic access (1)

- **Apache Nutch Agentic Access** — 24 operations · 10 acting

## Security (3)

- **Apache Nutch Authentication** — http · 1 scheme
- **Apache Nutch Domain Security** — TLSv1.3 · HSTS · DMARC
- **Apache Nutch Vulnerability Disclosure** — security.txt · contact published

## Plans (1)

- **Apache Nutch Plans Pricing**

## Use cases (6)

- **Enterprise Search** — Build enterprise search engines over internal or external web content using Nutch as the crawler and Solr/Elasticsearch as the search backend.
- **Research Data Collection** — Academic and research teams use Nutch for large-scale systematic web data collection and indexing.
- **Intranet Document Search** — Crawl and index intranet sites, wikis, and document repositories for internal enterprise search.
- **Web Archive Creation** — Create structured web archives compatible with CommonCrawl format for long-term data preservation.
- **SEO and Content Monitoring** — Monitor web content changes, track competitor sites, and analyze web structure at scale.
- **Custom Data Extraction Pipelines** — Build custom extraction pipelines using Nutch plugin architecture for targeted data acquisition tasks.

## Tags

Web Crawler, Indexing, Search, Apache, Java, Hadoop, Open-Source

---

Profiled by [API Evangelist](https://apievangelist.com) and published on [APIs.io](https://apis.io/providers/apache-nutch/). Scores are computed from the provider's own public artifacts under a published rubric.
