There are 4 repositories under tika topic.
Spark-Crawler: Apache Nutch-like crawler that runs on Apache Spark.
Code for Machine Learning with TensorFlow: 2nd Edition Published by Manning Publications
Viewers for statistics and dashboarding of Domain Search Engine data
Apache Tika bindings for PHP: extract text and metadata from documents, images and other formats
Tika-Similarity uses the Tika-Python package (Python port of Apache Tika) to compute file similarity based on Metadata features.
Interactive Image similarity and Visual Search and Retrieval application
ImageCat is an Apache OODT RADIX application that uses Apache Solr, Apache Tika and Apache OODT to ingest 10s of millions of files (images,but could be extended to other files) in place, and to extract metadata and OCR information from those files/images using Tika and Tesseract OCR.
Extract and Visualize location from any file
Geographic Place, Date/time, and Pattern entity extraction toolkit along with text extraction from unstructured data and GIS outputters.
Distributed, fault tolerant batch processing for Natural Language Applications and Search, using remote partitioning
Apache NiFi Custom Processor Extracting Text From Files with Apache Tika
Java web application taking IPFS hashes, extracting (textual) content and metadata through Apache's Tika.
A ruby wrapper for the Tika jar (tika-app.jar) that extracts text in a lot of formats from PDF, xls, doc, etc files
📄🚀 Unleash a powerful Document Search Engine with Apache NiFi for lightning-fast, comprehensive text indexing and search.
A suite of Machine Learning / Deep Learning Dockerfiles to allow Apache Tika to extract objects and to produce textual captions for images and video
Distill information about amendments to the Oregon Revised Statutes.
Extract text from a document by Apache Tika
Search Engine projects
Apache Tika Server as a Background Service in Node.js
An Elasticsearch engine plugin for Moodle's Global Search
A dataset downloaded from the deep and scientific web across three major Polar data centers for use in research.
qDesktopSearch - a Qt5 Desktop App for indexing & searching the files of the local machine
TYPO3 Extension: solr_file_indexer