Agent skill for pixel-grounded chart data extraction
-
Updated
Jul 3, 2026 - Python
Agent skill for pixel-grounded chart data extraction
Converters where figures survive. DOCX/XLSX to Markdown via native OOXML chart data: real numbers, OCR/VLM are only optional. CLI, MCP server, Docker image, and .mcpb bundle included.
doc-textify: offline, CPU-only PDF/image to Markdown & LLM-ready text converter. OCR with CJK normalization, layout recovery, two-column reading order, table/chart/formula extraction, RAG-ready chunking — no vision LLMs, no GPU, no cloud.
CUDA-accelerated PDF -> Markdown/HTML converter using Docling + IBM Granite Vision chart extraction
A complete end-to-end pipeline for extracting structured data from chart and graph images
Python library for extracting content from PowerPoint files including embedded charts and SmartArt. Built for RAG and document processing pipelines.
Extract charts, figures, and tables from academic PDFs for AI agent analysis
Add a description, image, and links to the chart-extraction topic page so that developers can more easily learn about it.
To associate your repository with the chart-extraction topic, visit your repo's landing page and select "manage topics."