Apache Tika Document Parser Agent · git:20260710.62afd0f · 2026-07-10 · sha256 8a950e5391c3653e
Apache Tika Document Parser Agent git:20260710.62afd0fA
Immutable. This exact content is served forever at /api/v1/blob/8a950e5391c3653e.
--- name: "Apache Tika Document Parser Agent" slug: "apache-tika-document-parser-agent" description: "Extracts text and metadata from 1000+ file formats using Apache Tika server REST API. Handles PDF OCR via Tesseract integration, Office document parsing, and email archive extraction with MIME detection." github_stars: 3703 verification: "security_reviewed" source: "https://github.com/apache/tika" author: "The Apache Software Foundation" category: "Data Extraction & Transformation" framework: "Gemini" tool_ecosystem: github_repo: "apache/tika" github_stars: 3703 --- # Apache Tika Document Parser Agent Extracts text and metadata from 1000+ file formats using Apache Tika server REST API. Handles PDF OCR via Tesseract integration, Office document parsing, and email archive extraction with MIME detection. ## Installation Requirements and caveats from upstream: - **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped. Basic usage or getting-started notes: - =========== - **Parse a file in Java:** - java - Source: https://github.com/apache/tika - Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md ## Source - [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-document-parser-agent/)