git:20260519.f531e55 to git:20260710.2925510

11 added, 14 removed. Audit A to B.

---
- name: "Apache Tika Content Extraction Hub"
- slug: "apache-tika-content-extraction-hub"
+ title: "Apache Tika Content Extraction Hub"
description: "Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection."
- github_stars: 3703
verification: "security_reviewed"
source: "https://github.com/apache/tika"
author: "The Apache Software Foundation"
- category: "Data Extraction & Transformation"
- framework: "Custom Agents"
+ category:
+ - "Data Extraction & Transformation"
+ framework:
+ - "Custom Agents"
tool_ecosystem:
github_repo: "apache/tika"
github_stars: 3703
---
# Apache Tika Content Extraction Hub
Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection.
## Installation
- Requirements and caveats from upstream:
- - **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.
-
- Basic usage or getting-started notes:
- - ===========
- - **Parse a file in Java:**
- - java
+ Choose whichever fits your setup:
- - Source: https://github.com/apache/tika
- - Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md
+ 1. Copy this skill folder into your local skills directory.
+ 2. Clone the repo and symlink or copy the skill into your agent workspace.
+ 3. Add the repo as a git submodule if you manage shared skills centrally.
+ 4. Install it through your internal provisioning or packaging workflow.
+ 5. Download the folder directly from GitHub and place it in your skills collection.
## Source
- [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-content-extraction-hub/)