Multimodal OCR
Nanonets / olmOCR / RolmOCR / Aya-Vision / Qwen2-VL-OCR
Nanonets / olmOCR / RolmOCR / Aya-Vision / Qwen2-VL-OCR
All in One Banana for you!
Detect AI-generated content in images, videos, and audio
Detect and annotate objects in images
Detect AI-generated images with ensemble and forensic tools
Generate any application by Vibe Coding it
Real-time video captioning powered by FastVLM
Embedded MinerU document extraction demo
A multilingual PDF translator that preserves document layout
Generate a podcast to discuss the topic of your choice!
Convert PDFs to text using OCR
Generate a preview image from a PDF file
Split & merge PDFs in-memory (fast and private!)
Extract text and metadata from PDF files
Generate spokenβready scripts from documents for podcasts, lectures, or summaries
Extract text from PDFs with OCR