A lot of public-record sites contain useful people data (phones, address history, relatives), but the data is locked inside messy HTML pages.
I experimented with building a pipeline that extracts those pages and converts them into structured fields automatically.
The interesting part wasn’t scraping — it was normalizing inconsistent formats across records.
Curious if anyone else here builds pipelines for turning messy web sources into structured datasets.
submitted by /u/Aggressive_Cut7433
[link] [comments]