Home Projects text-extract-api
text-extract-api

text-extract-api

by CatchTheTornado · GitHub

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

# anonymization# api# extract# json
View on GitHub
⭐ Stars
3.1k
🍴 Forks
277
🔥 Trending
+5 this week
📜 License
MIT
Commercial use OK
📅 Created
2024
🔄 Last commit
7 mo ago
🏷️ Category
anonymization
💻 Language
🖥️ Self-hostable
Likely
You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
text-extract-api — GitHub preview card
📈 Star history
3 1433 135
2026-07-072026-07-22
📄 About

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

Frequently asked questions

What is text-extract-api?

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

Is text-extract-api open source?

text-extract-api is an open-source project. It is released under the MIT license.

Is text-extract-api free?

Yes. text-extract-api is free and open source — you can use, modify and self-host it.

🏅 Maintainer of this project?
OpenSourceAI badge — text-extract-api

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![OpenSourceAI](https://opensourceai.tech/badge.php?tool=catchthetornado-text-extract-api)](https://opensourceai.tech/project/catchthetornado-text-extract-api.html)
More badge options →
🧬 Related projects🧬 View the DNA map →