apache-tika-content-extraction-hub · git:20260519.f531e55 · 2026-05-19 · sha256 0b54179c846399bc

apache-tika-content-extraction-hub git:20260519.f531e55A

Immutable. This exact content is served forever at /api/v1/blob/0b54179c846399bc.

---
name: "Apache Tika Content Extraction Hub"
slug: "apache-tika-content-extraction-hub"
description: "Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection."
github_stars: 3703
verification: "security_reviewed"
source: "https://github.com/apache/tika"
author: "The Apache Software Foundation"
category: "Data Extraction & Transformation"
framework: "Custom Agents"
tool_ecosystem:
  github_repo: "apache/tika"
  github_stars: 3703
---

# Apache Tika Content Extraction Hub

Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection.

## Installation

Requirements and caveats from upstream:
- **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.

Basic usage or getting-started notes:
- ===========
- **Parse a file in Java:**
- java

- Source: https://github.com/apache/tika
- Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-content-extraction-hub/)