I have processed the Epstein Files dataset and extracted 5,082 email threads with 16,447 individual messages. I used an LLM (xAI Grok 4.1 Fast via OpenRouter API) to parse the OCR’d text and extract structured email data.
Dataset available here: https://huggingface.co/datasets/notesbymuneeb/epstein-emails
submitted by /u/muneebdev
[link] [comments]