git:20260710.2925510 to git:20260710.62afd0f
14 added, 11 removed. Audit B to A.
---
- title: "Apache Tika Document Parser Agent"
+ name: "Apache Tika Document Parser Agent"
+ slug: "apache-tika-document-parser-agent"
description: "Extracts text and metadata from 1000+ file formats using Apache Tika server REST API. Handles PDF OCR via Tesseract integration, Office document parsing, and email archive extraction with MIME detection."
+ github_stars: 3703
verification: "security_reviewed"
source: "https://github.com/apache/tika"
author: "The Apache Software Foundation"
- category:
- - "Data Extraction & Transformation"
- framework:
- - "Gemini"
+ category: "Data Extraction & Transformation"
+ framework: "Gemini"
tool_ecosystem:
github_repo: "apache/tika"
github_stars: 3703
---
# Apache Tika Document Parser Agent
Extracts text and metadata from 1000+ file formats using Apache Tika server REST API. Handles PDF OCR via Tesseract integration, Office document parsing, and email archive extraction with MIME detection.
## Installation
- Choose whichever fits your setup:
+ Requirements and caveats from upstream:
+ - **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.
- 1. Copy this skill folder into your local skills directory.
- 2. Clone the repo and symlink or copy the skill into your agent workspace.
- 3. Add the repo as a git submodule if you manage shared skills centrally.
- 4. Install it through your internal provisioning or packaging workflow.
- 5. Download the folder directly from GitHub and place it in your skills collection.
+ Basic usage or getting-started notes:
+ - ===========
+ - **Parse a file in Java:**
+ - java
+
+ - Source: https://github.com/apache/tika
+ - Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md
## Source
- [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-document-parser-agent/)