apache-tika-document-extractor · git:20260518.2d3c27b · 2026-05-18 · sha256 becad01a817510a4
apache-tika-document-extractor git:20260518.2d3c27bA
Immutable. This exact content is served forever at /api/v1/blob/becad01a817510a4.
--- name: "Apache Tika Document Extractor" slug: "apache-tika-document-extractor" description: "Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode." github_stars: 3695 verification: "listed" source: "https://github.com/apache/tika" category: "Data Extraction & Transformation" framework: "Codex" tool_ecosystem: github_repo: "apache/tika" github_stars: 3695 --- # Apache Tika Document Extractor Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode. ## Installation Requirements and caveats from upstream: - **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped. Basic usage or getting-started notes: - =========== - **Parse a file in Java:** - java - Source: https://github.com/apache/tika - Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md ## Source - [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-document-extractor/)