apache-tika-document-extractor ยท diff
git:20260519.f531e55 to git:20260710.2925510
11 added, 14 removed. Audit A to B.
---
- name: "Apache Tika Document Extractor"
- slug: "apache-tika-document-extractor"
+ title: "Apache Tika Document Extractor"
description: "Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode."
- github_stars: 3695
verification: "security_reviewed"
source: "https://github.com/apache/tika"
- category: "Data Extraction & Transformation"
- framework: "Codex"
+ category:
+ - "Data Extraction & Transformation"
+ framework:
+ - "Codex"
tool_ecosystem:
github_repo: "apache/tika"
github_stars: 3695
---
# Apache Tika Document Extractor
Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode.
## Installation
- Requirements and caveats from upstream:
- - **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.
-
- Basic usage or getting-started notes:
- - ===========
- - **Parse a file in Java:**
- - java
+ Choose whichever fits your setup:
- - Source: https://github.com/apache/tika
- - Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md
+ 1. Copy this skill folder into your local skills directory.
+ 2. Clone the repo and symlink or copy the skill into your agent workspace.
+ 3. Add the repo as a git submodule if you manage shared skills centrally.
+ 4. Install it through your internal provisioning or packaging workflow.
+ 5. Download the folder directly from GitHub and place it in your skills collection.
## Source
- [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-document-extractor/)