{"id":41920,"date":"2026-07-30T05:27:13","date_gmt":"2026-07-30T03:27:13","guid":{"rendered":"https:\/\/www.graviton.at\/letterswaplibrary\/self-promotion-i-built-a-public-dataset-from-21237-pages-of-declassified-mkultra-and-related-docs-and-put-it-on-hugging-face\/"},"modified":"2026-07-30T05:27:13","modified_gmt":"2026-07-30T03:27:13","slug":"self-promotion-i-built-a-public-dataset-from-21237-pages-of-declassified-mkultra-and-related-docs-and-put-it-on-hugging-face","status":"publish","type":"post","link":"https:\/\/www.graviton.at\/letterswaplibrary\/self-promotion-i-built-a-public-dataset-from-21237-pages-of-declassified-mkultra-and-related-docs-and-put-it-on-hugging-face\/","title":{"rendered":"[self-promotion] I Built A Public Dataset From 21,237 Pages Of Declassified MKULTRA And Related Docs And Put It On Hugging Face"},"content":{"rendered":"<p><!-- SC_OFF --><\/p>\n<div class=\"md\">\n<p>Until recently, the surviving historical records from the CIA&#8217;s MKULTRA and related programs were very difficult to search and analyze. So I ran 21,237 document page images through MinerU OCR to generate clean text transcripts, then produced redaction mappings to go with every page transcript. Original page images are stored on IPFS and are available for public download. The dataset is <a href=\"https:\/\/huggingface.co\/datasets\/peers-ai\/mkultra-foia-mineru-archive\">available on Hugging Face here<\/a>.<\/p>\n<\/div>\n<p><!-- SC_ON -->   submitted by   <a href=\"https:\/\/www.reddit.com\/user\/fixingbrokenrobots\"> \/u\/fixingbrokenrobots <\/a> <br \/> <span><a href=\"https:\/\/www.reddit.com\/r\/datasets\/comments\/1vaho5q\/selfpromotion_i_built_a_public_dataset_from_21237\/\">[link]<\/a><\/span>   <span><a href=\"https:\/\/www.reddit.com\/r\/datasets\/comments\/1vaho5q\/selfpromotion_i_built_a_public_dataset_from_21237\/\">[comments]<\/a><\/span><\/p><div class='watch-action'><div class='watch-position align-right'><div class='action-like'><a class='lbg-style1 like-41920 jlk' href='javascript:void(0)' data-task='like' data-post_id='41920' data-nonce='3bde7dc5cf' rel='nofollow'><img class='wti-pixel' src='https:\/\/www.graviton.at\/letterswaplibrary\/wp-content\/plugins\/wti-like-post\/images\/pixel.gif' title='Like' \/><span class='lc-41920 lc'>0<\/span><\/a><\/div><\/div> <div class='status-41920 status align-right'><\/div><\/div><div class='wti-clear'><\/div>","protected":false},"excerpt":{"rendered":"<p>Until recently, the surviving historical records from the CIA&#8217;s MKULTRA and related programs were very difficult to&#8230;<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[85],"tags":[],"class_list":["post-41920","post","type-post","status-publish","format-standard","hentry","category-datatards","wpcat-85-id"],"_links":{"self":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts\/41920","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/comments?post=41920"}],"version-history":[{"count":0,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts\/41920\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/media?parent=41920"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/categories?post=41920"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/tags?post=41920"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}