Olmocr
A toolkit designed for converting and linearizing PDFs to create datasets optimized for large language model (LLM) training and evaluation.
Olmocr is built in Python, distributed under the Apache License 2.0, 17.1k GitHub stars, latest release v0.4.27.
When to use Olmocr
Olmocr is listed here as a AI project. The directory calls out PDF linearization for dataset creation, Optimized for large language model (LLM) workflows, Supports automated text extraction as capabilities associated with it.
Other recorded traits for Olmocr include Facilitates preparation of training datasets, Command-line utility for easy usage.
Besides AI, this page also files Olmocr under Tools, Development, Utils.
Olmocr compared with
Records in this directory name pdfplumber, pdftotext, pdfminer.six, unstructured as products people compare with Olmocr. That list is editorial metadata, not a claim that Olmocr replaces each of them.
What the Olmocr stats reflect
GitHub currently shows 17.1k GitHub stars, 1.4k forks, 74 open issues, latest tracked release v0.4.27. Star and activity counts here are a snapshot used as a proxy for community adoption, not a quality score.
Stats refreshed
- Language
- Python
- Latest Release
- v0.4.27
- License
- Apache License 2.0
Our Newsletter
Get new AI tools right in your inbox
Get short emails with useful ai projects, releases, and repos worth watching.
Key features of Olmocr
- PDF linearization for dataset creation
- Optimized for large language model (LLM) workflows
- Supports automated text extraction
- Facilitates preparation of training datasets
- Command-line utility for easy usage
Recorded alternatives to Olmocr
Olmocr resources
Olmocr on GitHub
Frequently asked questions
What is Olmocr?
A toolkit designed for converting and linearizing PDFs to create datasets optimized for large language model (LLM) training and evaluation. This directory highlights PDF linearization for dataset creation, Optimized for large language model (LLM) workflows, Supports automated text extraction.
Is Olmocr free to use?
Olmocr is published as open source under the Apache License 2.0. The directory lists PDF linearization for dataset creation, Optimized for large language model (LLM) workflows, Supports automated text extraction among its recorded capabilities.
What language is Olmocr written in, and what is the latest release?
Olmocr is written primarily in Python. The latest release tracked on this page is v0.4.27.
How widely is Olmocr used on GitHub?
Olmocr has about 17.1k GitHub stars. It also has about 1.4k forks. Those counts are a snapshot of community attention, not a ranking of quality.
What do people compare Olmocr with?
This directory records pdfplumber, pdftotext, pdfminer.six, unstructured as comparison points for Olmocr.
Related tools
Unstructured
Unstructured is an open-source ETL solution that converts complex documents into structured data for language models, featuring enterprise-grade capabilities like workflow orchestration, document partitioning, enrichment, chunking, and embedding.
Pdfplumber
pdfplumber is a powerful Python library for extracting detailed information from PDF files, including text, tables, individual characters, shapes, and more. It allows precise programmatic access to underlying PDF content, making it highly useful for data extraction, automation, and analysis tasks.
Cognee
Memory for AI Agents in 5 lines of code. Cognee provides a lightweight and efficient memory management solution tailored for AI agent development.
Cua
Open-source infrastructure for training, evaluating, and benchmarking AI-powered Computer-Use Agents that control full desktop environments (macOS, Linux, Windows) with integrated sandboxes, SDKs, and reproducible evaluation tools.
Developer
The pioneering library allowing the embedding of a developer agent into your application, enhancing interactivity and automation.
Composio
Equips your AI agents & LLMs with over 100 high-quality integrations via function calling.