{"id":41783,"date":"2026-07-21T22:25:52","date_gmt":"2026-07-21T20:25:52","guid":{"rendered":"https:\/\/www.graviton.at\/letterswaplibrary\/getting-30-years-of-sec-company-fundamentals-is-hard\/"},"modified":"2026-07-21T22:25:52","modified_gmt":"2026-07-21T20:25:52","slug":"getting-30-years-of-sec-company-fundamentals-is-hard","status":"publish","type":"post","link":"https:\/\/www.graviton.at\/letterswaplibrary\/getting-30-years-of-sec-company-fundamentals-is-hard\/","title":{"rendered":"Getting 30+ Years Of SEC Company Fundamentals Is HARD!"},"content":{"rendered":"<p><!-- SC_OFF --><\/p>\n<div class=\"md\">\n<p>I&#8217;m the founder of StockFit API &#8211; SEC sourced clean fundamentals for all US Companies (delisted or not) &#8211; all Point-In-Time data perfect for back testing.<\/p>\n<p>One of my subscribers has pushed me to investigate what it would take to widen the coverage to the pre-XBRL time.<\/p>\n<p><strong>If you don&#8217;t know what that means:<\/strong><br \/> In 2009, XBRL was mandated my the SEC, which from that time on provides a structured data format to extract the actual fundamentals a company reported.<\/p>\n<p>Before that time, there where 2 distinct data format time frames that you will have to prepare for if you want that data too:<br \/> 1. <strong>Pre 2001: FDS (Financial Data Schedule)<\/strong> &#8211; a semi-structured way data was reported, similar to XML<br \/> 2. <strong>2001 &#8211; 2009: No structure at all<\/strong> &#8211; a full mix of HTML\/ASCII text\/tables <\/p>\n<p>FDS is kinda ok to parse, but comes with many exceptions across filers. You can easily get to a coverage of 90%+<\/p>\n<p>The nightmare starts in 2001. There really is no structural mandate whatsoever that you can rely on! Everything has to be tested empirically and your parser needs to be able to handle everything, validate as much as it can, and provide an observation layer for you to ensure data integrity.<\/p>\n<p>This is probably the most challenging piece in my entire stack to get right. But in the end, I would be able to claim that I serve 30+ years of historical fundamentals, and that is absolutely worth it.<\/p>\n<p>This effort is NOT finished yet. Until now, I&#8217;m serving 20+ years of fundamentals that are of very high quality\/accuracy. Getting that next batch to that same bar is something I&#8217;m trying to get to now.<\/p>\n<p>But I&#8217;ll absolutely take one step at a time to get there. Otherwise this will get ugly very quickly. And there is nothing worse than serving wrong data.<\/p>\n<p>I&#8217;m curious if anyone has done this pre-XBRL parsing before? Any lessons you can share after having gone through that? <\/p>\n<p>And from the consumer side: How much do you care in your use case about fundamentals that are pre-2009? <\/p>\n<\/div>\n<p><!-- SC_ON -->   submitted by   <a href=\"https:\/\/www.reddit.com\/user\/Either_Door_5500\"> \/u\/Either_Door_5500 <\/a> <br \/> <span><a href=\"https:\/\/www.reddit.com\/r\/datasets\/comments\/1v2sp25\/getting_30_years_of_sec_company_fundamentals_is\/\">[link]<\/a><\/span>   <span><a href=\"https:\/\/www.reddit.com\/r\/datasets\/comments\/1v2sp25\/getting_30_years_of_sec_company_fundamentals_is\/\">[comments]<\/a><\/span><\/p><div class='watch-action'><div class='watch-position align-right'><div class='action-like'><a class='lbg-style1 like-41783 jlk' href='javascript:void(0)' data-task='like' data-post_id='41783' data-nonce='ee32349510' rel='nofollow'><img class='wti-pixel' src='https:\/\/www.graviton.at\/letterswaplibrary\/wp-content\/plugins\/wti-like-post\/images\/pixel.gif' title='Like' \/><span class='lc-41783 lc'>0<\/span><\/a><\/div><\/div> <div class='status-41783 status align-right'><\/div><\/div><div class='wti-clear'><\/div>","protected":false},"excerpt":{"rendered":"<p>I&#8217;m the founder of StockFit API &#8211; SEC sourced clean fundamentals for all US Companies (delisted or&#8230;<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[85],"tags":[],"class_list":["post-41783","post","type-post","status-publish","format-standard","hentry","category-datatards","wpcat-85-id"],"_links":{"self":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts\/41783","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/comments?post=41783"}],"version-history":[{"count":0,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts\/41783\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/media?parent=41783"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/categories?post=41783"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/tags?post=41783"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}