17.1k

Olmocr

A toolkit designed for converting and linearizing PDFs to create datasets optimized for large language model (LLM) training and evaluation.

Olmocr is built in Python, distributed under the Apache License 2.0, 17.1k GitHub stars, latest release v0.4.27.

When to use Olmocr

Olmocr is listed here as a AI project. The directory calls out PDF linearization for dataset creation, Optimized for large language model (LLM) workflows, Supports automated text extraction as capabilities associated with it.

Other recorded traits for Olmocr include Facilitates preparation of training datasets, Command-line utility for easy usage.

Besides AI, this page also files Olmocr under Tools, Development, Utils.

Olmocr compared with

Records in this directory name pdfplumber, pdftotext, pdfminer.six, unstructured as products people compare with Olmocr. That list is editorial metadata, not a claim that Olmocr replaces each of them.

What the Olmocr stats reflect

GitHub currently shows 17.1k GitHub stars, 1.4k forks, 74 open issues, latest tracked release v0.4.27. Star and activity counts here are a snapshot used as a proxy for community adoption, not a quality score.

Stats refreshed

Language
Python
Latest Release
v0.4.27
License
Apache License 2.0

Our Newsletter

Get new AI tools right in your inbox

Get short emails with useful ai projects, releases, and repos worth watching.


Key features of Olmocr

  • PDF linearization for dataset creation
  • Optimized for large language model (LLM) workflows
  • Supports automated text extraction
  • Facilitates preparation of training datasets
  • Command-line utility for easy usage

Recorded alternatives to Olmocr


Olmocr resources


Olmocr on GitHub

Stars
17.1k
Open Issues
74
Forks
1.4k

Frequently asked questions

What is Olmocr?

A toolkit designed for converting and linearizing PDFs to create datasets optimized for large language model (LLM) training and evaluation. This directory highlights PDF linearization for dataset creation, Optimized for large language model (LLM) workflows, Supports automated text extraction.

Is Olmocr free to use?

Olmocr is published as open source under the Apache License 2.0. The directory lists PDF linearization for dataset creation, Optimized for large language model (LLM) workflows, Supports automated text extraction among its recorded capabilities.

What language is Olmocr written in, and what is the latest release?

Olmocr is written primarily in Python. The latest release tracked on this page is v0.4.27.

How widely is Olmocr used on GitHub?

Olmocr has about 17.1k GitHub stars. It also has about 1.4k forks. Those counts are a snapshot of community attention, not a ranking of quality.

What do people compare Olmocr with?

This directory records pdfplumber, pdftotext, pdfminer.six, unstructured as comparison points for Olmocr.