apache-tika-content-extraction-hub · git:20260710.2925510 · 2026-07-10 · sha256 5b825d4bbdaa6668
apache-tika-content-extraction-hub git:20260710.2925510B
Immutable. This exact content is served forever at /api/v1/blob/5b825d4bbdaa6668.
--- title: "Apache Tika Content Extraction Hub" description: "Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection." verification: "security_reviewed" source: "https://github.com/apache/tika" author: "The Apache Software Foundation" category: - "Data Extraction & Transformation" framework: - "Custom Agents" tool_ecosystem: github_repo: "apache/tika" github_stars: 3703 --- # Apache Tika Content Extraction Hub Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection. ## Installation Choose whichever fits your setup: 1. Copy this skill folder into your local skills directory. 2. Clone the repo and symlink or copy the skill into your agent workspace. 3. Add the repo as a git submodule if you manage shared skills centrally. 4. Install it through your internal provisioning or packaging workflow. 5. Download the folder directly from GitHub and place it in your skills collection. ## Source - [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-content-extraction-hub/)