News
AI Summary
10 Aug 202627 Safar 1448 AH
Old OCR text cripples language model training, and FineBooks wants to fix that at scale

Old OCR text cripples language model training, and FineBooks wants to fix that at scale

The FineBooks project from Hugging Face and EleutherAI tested 14 open-source OCR models on over 2,000 historical book pages. The top model, dots.mocr, achieved a character accuracy of 97.6% at a cost of under two dollars per thousand pages. While this accuracy is sufficient for AI training data, it is not yet adequate for scholarly transcriptions, according to the team.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In