[self-promotion] I Built A Public Dataset From 21,237 Pages Of Declassified MKULTRA And Related Docs And Put It On Hugging Face

Until recently, the surviving historical records from the CIA’s MKULTRA and related programs were very difficult to search and analyze. So I ran 21,237 document page images through MinerU OCR to generate clean text transcripts, then produced redaction mappings to go with every page transcript. Original page images are stored on IPFS and are available for public download. The dataset is available on Hugging Face here.

submitted by /u/fixingbrokenrobots
[link] [comments]

Leave a Reply

Your email address will not be published. Required fields are marked *