Hey, if someone is looking for a large dataset of OCRed (various quality) text content in different languages, mostly for LLM training, feel free to reach me (I’m the maintainer) here or at the site. There you also may find a demo for testing quality of the data.
submitted by /u/Infinite-Band6504
[link] [comments]