{"id":42578,"date":"2026-09-12T17:27:13","date_gmt":"2026-09-12T15:27:13","guid":{"rendered":"https:\/\/www.graviton.at\/letterswaplibrary\/dataset-pinga-fogo-with-chico-xavier-brazilian-tv-1971-345-speech-turns-from-6h-of-live-interview-audio-portuguese-asr-timestamps\/"},"modified":"2026-09-12T17:27:13","modified_gmt":"2026-09-12T15:27:13","slug":"dataset-pinga-fogo-with-chico-xavier-brazilian-tv-1971-345-speech-turns-from-6h-of-live-interview-audio-portuguese-asr-timestamps","status":"publish","type":"post","link":"https:\/\/www.graviton.at\/letterswaplibrary\/dataset-pinga-fogo-with-chico-xavier-brazilian-tv-1971-345-speech-turns-from-6h-of-live-interview-audio-portuguese-asr-timestamps\/","title":{"rendered":"[Dataset] Pinga-Fogo With Chico Xavier (Brazilian TV, 1971) \u2014 345 Speech Turns From ~6h Of Live Interview Audio, Portuguese, ASR + Timestamps"},"content":{"rendered":"<p><!-- SC_OFF --><\/p>\n<div class=\"md\">\n<p>I\u2019ve published a structured transcript dataset of <em>Pinga-Fogo<\/em>, two live TV interviews with Brazilian medium Chico Xavier, broadcast by TV Tupi in 1971.<\/p>\n<p>Together, the two programs amount to roughly 6 hours of material and are among the longest surviving recordings of Chico Xavier answering questions live before a panel of journalists.<\/p>\n<p>Dataset:<\/p>\n<p><a href=\"https:\/\/huggingface.co\/datasets\/ia-espirita\/pinga-fogo-chico-xavier\">https:\/\/huggingface.co\/datasets\/ia-espirita\/pinga-fogo-chico-xavier<\/a><\/p>\n<p><strong>What\u2019s in it<\/strong><\/p>\n<p>345 speech turns, including 115 answers by Chico Xavier.<\/p>\n<p>Available in JSONL, Parquet, and combined CSV. Each turn includes speaker, turn type, topic label, timestamps, transcript text, source audio, and review flags.<\/p>\n<p>Portuguese only.<\/p>\n<p><strong>How it was made<\/strong><\/p>\n<p>Internet Archive audio \u2192 Whisper large-v3 \u2192 LLM-assisted segmentation and labeling.<\/p>\n<p>The LLM did <strong>not<\/strong> rewrite the transcript. It only worked with segment indices, boundaries, speaker\/turn classification, and topic labels.<\/p>\n<p>Every word in the transcript comes directly from Whisper output.<\/p>\n<p><strong>Known limitations<\/strong><\/p>\n<p>This is still unreviewed ASR from degraded 1971 recordings, so there are transcription errors, especially in proper names and numbers.<\/p>\n<p>Every record is marked:<\/p>\n<p><code>revisado_por_humano: false<\/code><\/p>\n<p>If you want to quote something, check the original audio first. Timestamps are included for that purpose.<\/p>\n<p>The date of the second program is inconsistent across historical sources, so I recorded only December 1971 rather than forcing an exact day.<\/p>\n<p>Speaker attribution is also left as <code>desconhecido<\/code> whenever I couldn&#8217;t identify someone confidently.<\/p>\n<p><strong>Why I made it<\/strong><\/p>\n<p>I\u2019m building a retrieval system over historical Spiritist literature and primary sources, and I couldn\u2019t find a structured, timestamped version of <em>Pinga-Fogo<\/em> anywhere.<\/p>\n<p>Potential uses include pt-BR ASR benchmarking, speaker-turn segmentation, diarization, information retrieval, and long-form QA.<\/p>\n<p>Human review is the obvious next step.<\/p>\n<p>Corrections and PRs are very welcome.<\/p>\n<\/div>\n<p><!-- SC_ON -->   submitted by   <a href=\"https:\/\/www.reddit.com\/user\/SideSuspicious8083\"> \/u\/SideSuspicious8083 <\/a> <br \/> <span><a href=\"https:\/\/www.reddit.com\/r\/datasets\/comments\/1weebwn\/dataset_pingafogo_with_chico_xavier_brazilian_tv\/\">[link]<\/a><\/span>   <span><a href=\"https:\/\/www.reddit.com\/r\/datasets\/comments\/1weebwn\/dataset_pingafogo_with_chico_xavier_brazilian_tv\/\">[comments]<\/a><\/span><\/p><div class='watch-action'><div class='watch-position align-right'><div class='action-like'><a class='lbg-style1 like-42578 jlk' href='javascript:void(0)' data-task='like' data-post_id='42578' data-nonce='4ef4545ee1' rel='nofollow'><img class='wti-pixel' src='https:\/\/www.graviton.at\/letterswaplibrary\/wp-content\/plugins\/wti-like-post\/images\/pixel.gif' title='Like' \/><span class='lc-42578 lc'>0<\/span><\/a><\/div><\/div> <div class='status-42578 status align-right'><\/div><\/div><div class='wti-clear'><\/div>","protected":false},"excerpt":{"rendered":"<p>I\u2019ve published a structured transcript dataset of Pinga-Fogo, two live TV interviews with Brazilian medium Chico Xavier,&#8230;<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[85],"tags":[],"class_list":["post-42578","post","type-post","status-publish","format-standard","hentry","category-datatards","wpcat-85-id"],"_links":{"self":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts\/42578","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/comments?post=42578"}],"version-history":[{"count":0,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts\/42578\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/media?parent=42578"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/categories?post=42578"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/tags?post=42578"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}