The FineBooks project, a collaboration involving Hugging Face and EleutherAI, recently tested 14 open-source OCR models on over 2,000 pages of historical books to improve text recognition accuracy for AI training datasets, according to The Decoder.

The leading model, dots.mocr, achieved a character accuracy rate of 97.6%, with processing costs under two dollars per thousand pages. While this level of accuracy is sufficient for training AI, it still falls short of the quality required for scholarly transcription work.

This development is significant for Japanese markets where digitization of historical documents and efficient AI training are key to advancing language technology and data services in finance and equities sectors.