Category: Datatards

Here you can observe the biggest nerds in the world in their natural habitat, longing for data sets. Not that it isn’t interesting, i’m interested. Maybe they know where the chix are. But what do they need it for? World domination?

[Self-promotion] Free Entity Lookup: 521M Legal Entities Across 309 Jurisdictions, No Account Needed To Search

Disclosing this up front – I’m at Veridion.

We built this because one of our leads mentioned that they pay a registry data provider six figures a year. For legal names, identifiers and registered addresses. We already had that data. Not as a side project, it’s the foundation layer under our enterprise product, and the pipelines were already running. So opening it up cost us close to nothing, which is sort of the point: the collection isn’t what makes it expensive elsewhere.

registry-lookup.com – 521M legal entities, 309 jurisdictions, 244 countries. Search on the site is free with no account. There’s an API at 5,000 calls a month if you want it programmatically, that one needs a work email.

What you get per entity: legal name, registry number, jurisdiction code, status, incorporation date, legal form, registered address, and identifiers like tax IDs, VAT where the registry publishes them.

So it’s an enumeration and triage tool. It answers “does this entity exist, what’s its number, is it active, where is it registered.”

Which jurisdictions do you currently have no good way to check?

submitted by /u/OkHeat6599
[link] [comments]

Seeking Anonymized Field Data Collection Datasets For An Open Benchmark

#

Hi everyone,

I’m working on an initiative to create an **open benchmark dataset for field data quality assurance**.

Today, there are many excellent digital data collection platforms—such as KoboToolbox, SurveyCTO, ODK, CommCare, Survey Solutions, CSPro, and others—but there are very few publicly available datasets that developers and researchers can use to evaluate field data quality tools.

I’m looking for individuals or organizations that may be willing to share **completed, fully anonymized datasets** from field data collection projects, where they have the necessary permissions to do so.

I’m especially interested in datasets that include:

* GPS coordinates (or generalized locations)
* Interview photos
* Audio recordings
* Interview start and end times
* Submission timestamps
* Enumerator IDs (anonymized)
* Supervisor review outcomes or quality flags (if available)

These datasets will help create a community benchmark for testing quality assurance methods such as:

* GPS verification
* Duplicate image detection
* Audio quality assessment
* Interview duration analysis
* Duplicate submission detection
* Fieldwork anomaly detection

The objective is to create a resource that benefits researchers, NGOs, software developers, and the wider field data collection community by making it easier to evaluate and improve quality assurance tools.

If your organization has a completed project that could be shared in an anonymized form—or if you know of existing public datasets—I would greatly appreciate hearing from you.

I’m also happy to discuss data-sharing agreements, attribution, licensing, or any requirements needed to ensure the data is used responsibly.

Thank you!

submitted by /u/Significant-West-492
[link] [comments]

Trying To Learn How To Use API To Extract Data

Hello! I’m a complete newbie in Data Science and I’m trying to learn how to get data from an API. I understand an API could be public or could require authentication.

I worked with CVS files and I wanted to experience or practice getting data from APIs.

I’m getting familiar with Python so I was wondering if you could help me with the following issues:

  1. Trying to understand and practice the different methods you can use API to request data (I am not sure if it has to be from a Dataset formar or can it be any kind of format) with Python

  2. What are some good options to get APIs to work on data Science

  3. I am not even close to get to a point where I am able to do Reproducible projects/models but I do wonder how including an API (understanding that it is some kind of “personal Key”) to share my code and people to be able to use it.

Hope I made sense of what my doubts are and I apologize in advance if I seem confused about some terms (I do think I am).

submitted by /u/MyNameCouldBeMarion
[link] [comments]

[Dataset] Dubai Residential Sale Prices And Volumes, Monthly January 2008 To July 2026, From Land Department Transactions

What: monthly citywide residential median AED per square foot, a 5-month centred average, an index rebased to 100 at January 2008, and monthly sales counts. 1,080,194 transactions across 223 months. A matching series for registered leases runs from May 2010.

Source: Dubai Land Department transaction and lease records, which are public.

Repo, with both series, method and licence: https://github.com/dataHabibi/dubai-price-index

Columns:

  • month
  • sales_count
  • median_aed_per_sqft
  • ma5_aed_per_sqft
  • index_base100
  • provisional

Two things to know before you use it.

The provisional column marks the last two months, where the centred average still has fewer than two later months to work with. Their raw median and sales count are fine, it is the smoothed value and the index that will keep moving.

Sales counts for recent months are understated. Registrations land one to two months after the deal closes, so the tail of that column is still filling in. Do not read the recent drop as a fall in demand.

What it is not: a repeat sales or hedonic index. It is a median, so it is not quality adjusted. Shifts in what sells, off plan against ready, apartment against villa, which communities are active, move this line without any individual property changing price. Treat it as a market thermometer.

CC BY 4.0. Refreshed monthly by a scheduled job, so the committed files track the live series.

submitted by /u/datahabibi
[link] [comments]

[self-promotion]Python Developer Available For Web Scraping & Automation Projects

Freelance Python developer available for projects involving web scraping and automation.

Skills:

Web scraping (Scrapy, Selenium, Playwright, BeautifulSoup)

Python automation scripts

API development and integration

Data extraction and ETL pipelines

FastAPI and Flask

Browser automation

CSV, Excel, JSON, and database processing

Docker and Linux deployment

Past work:

Lead generation scrapers

Google Maps data extraction

Business automation tools

Custom APIs and data pipelines

Open to one-time projects and long-term collaborations.

DM me if you need help automating a workflow or collecting data.

submitted by /u/HackerThing
[link] [comments]

[self-promotion] I Built A Public Dataset From 21,237 Pages Of Declassified MKULTRA And Related Docs And Put It On Hugging Face

Until recently, the surviving historical records from the CIA’s MKULTRA and related programs were very difficult to search and analyze. So I ran 21,237 document page images through MinerU OCR to generate clean text transcripts, then produced redaction mappings to go with every page transcript. Original page images are stored on IPFS and are available for public download. The dataset is available on Hugging Face here.

submitted by /u/fixingbrokenrobots
[link] [comments]

What Are The Best Publicly Available “uncensored” Datasets?

I use “Heretic” library on models to liberate them from their safeguards, but while checking their “uncensoredness”, I found they can hallucinate a lot. You know, it’s basically like a child who’s now allowed to use the F word once and he says “Fred” instead of the actual thing.

So I think if the models train on valid uncensored data (specially if they start Grokking) the results can improve. So I am using for these types of datasets to test my theory.

submitted by /u/Haghiri75
[link] [comments]

¿Does Anyone Know Where Can I Sell A Dataset With 10,000 Chines-related Classified News?

I’ve been working on a dataset for a research project on how China is portrayed in the media. It currently contains just over 10,000 news articles from both Chinese and Western news outlets.

Each article is classified by topic and by the way China is portrayed (e.g. positive, negative, threat, Xi-centered, neutral, etc.). The dataset was originally created for academic research, but I’m now wondering whether it could also have commercial value.

I’m not trying to sell it here, just looking for advice. Has anyone here ever licensed or sold a specialized dataset like this? Who would actually be interested in buying it? AI companies, media intelligence firms, universities, think tanks…? Or are datasets like this generally expected to be open source?

I’d really appreciate hearing from anyone who has experience commercializing niche datasets or knows how this market works.

submitted by /u/No_Programmer_2947
[link] [comments]

[Showcase] 54 High-value Real-time JSON Datasets For LLMs And Developers

Hey builders! 🤖

I’ve launched a project called x402 Data Hub (https://x402datahub.io). It is designed to solve a major problem for autonomous AI agents: accessing premium real-time data without expensive monthly subscriptions.

We offer 54 structured JSON datasets across:

– SaaS & AI API Pricing (LLM cost comparison, cloud GPU indexes)

– Live multi-chain Gas metrics (Base/Arbitrum)

– Tech jobs & developer market rates

– Expat living data (rent index, digital nomad visas, tax rules)

Zero API keys or email registrations required. Access feeds instantly using HTTP 402 pay-per-request protocol ($0.01 USDC on Base or Arbitrum).

🎁 We offer 50 free daily requests per dataset for developer testing.

Read the agent-friendly specification: https://x402datahub.io/llms.txt

Let me know what datasets you would like to see added next!

submitted by /u/x402DataHub
[link] [comments]

How Are You Guys Handling Financial Disclosures & Unstructured Data For Chinese (A-shares) And HK Stock Markets?

Hi all,

I’ve been working on a financial research project that involves analyzing company filings and disclosures for A-shares (Shanghai/Shenzhen) and HKEx listed entities.

Coming from a Western market background, the biggest pain points I’ve noticed are the language barrier, disparate filing locations, and the lack of structured APIs formatted for LLMs/RAG.

For those who cover APEx or emerging markets:

  1. What tools or data providers are you currently using for CN/HK stock filings?
  2. How do you handle language translation and footnote extraction in your pipeline?

Would love to exchange ideas with anyone working on similar Asia-Pacific equity pipelines!

submitted by /u/GapLucky1794
[link] [comments]

I Built An API That Turns Messy SEC Filings (insider Trades, 13F, Activist Stakes) Into Clean JSON. The Month-one Reality Was Ugly.

Quick background: public companies have to file everything with the SEC. Who’s buying their own stock, what hedge funds hold, who just crossed 5% with an activist stake. It’s all public and free on sec.gov. It’s also served as XML that makes you want to quit programming, and every filer formats it differently.

Edgrapi does that parsing once and hands it back as clean JSON. One call, one URL. That’s the product.

Month one taught me things I didn’t expect.

The data lies if you read it literally. Michael Burry’s latest fund filing is 66% Palantir and Nvidia. Except they’re puts, bearish bets. A naive parser reports him as massively long two stocks he’s actually betting against. Half the fund-tracker tools out there get this wrong, and getting it right turned out to be most of the value.

Postgres quietly broke every login for a week. Worked perfectly on my laptop, failed in production. A floating-point column was rounding my 10-digit timestamps, so every login token was born already expired by about three hours. SQLite stored it fine, Postgres didn’t. I only caught it because I logged the raw stored value instead of the error message.

Every API call took 2.6 seconds, no matter what. Cheap call, heavy call, same 2.6s. That flatness was the tell: it wasn’t the work, it was nine database round-trips to a server in another region. Collapsed it to two. It’s 0.7s now.

submitted by /u/Capedcrusader1923
[link] [comments]

Best Open-source Clean Speech And Ambient Noise Datasets For Training An Edge AI Audio Denoiser?

We are building an edge-AI audio noise-reduction system on an ESP32-S3.

Our architecture uses a lightweight GRUNet (~59k parameters) to output a dynamic gain mask on a 44-band Mel-spectrogram.

​I need gigabytes of audio to train the model. Does anyone have recommendations for the best open-source datasets for:

1> ​Clean, isolated human speech.

2> ​Diverse ambient background noise (traffic, crowds, machinery, etc.).

​Also, any tips or open-source scripts for artificially mixing these at different Signal-to-Noise Ratios (SNRs) before generating the 16kHz Mel-spectrograms would be hugely appreciated!

submitted by /u/saikat_munshib
[link] [comments]

[self-promotion] [PAID] Podcast Sponsorship Dataset: Which Brands Sponsor Which Shows, With The Verbatim Evidence Line For Every Record (free Tier Available)

Disclosure: I built this, it’s my project, and paid tiers exist. There’s a free tier and everything shown below is viewable without signing up.

What it is: structured sponsorship records extracted from public podcast RSS show notes. One row per (brand, episode):

brand (canonically resolved) | show | episode | publish date | promo code | promo URL + registrable domain | sponsor type (paid / affiliate / house ad) | confidence | confidence tier | first seen | last seen | the verbatim sentence the claim came from

Sample rows straight out of the DB:

– AG1 on Huberman Lab, 2026-07-27, evidence: “AG1: https://drinkag1.com/huberman

– Visible on Good Hang with Amy Poehler, code HANG, 2 episodes, 14-day span

– Saily on Machtwechsel (German news podcast), code “Machtwechsel”, 3 episodes over 18 days

Method, since this sub cares about it: LLM extraction over the show-notes text, then a deterministic brand-resolution layer on top. Domain evidence merges entities first (drinkag1.com and athleticgreens.com collapse into one AG1 entity), exact normalized-name match second, and anything that is merely name-similar goes to an adjudication queue and is never auto-merged. That last rule is what keeps Dove the soap separate from Dove the chocolate. Every record retains its source sentence so any claim can be audited by hand.

Honest limits, up front:

– Show notes only. Ads that exist purely in audio and never appear in the notes are invisible to this. Transcript coverage is not built yet.

– The corpus is small right now: 314 episodes across 93 shows, US + DE + FR. It grows daily but this is not a historical archive.

– I am deliberately not publishing an accuracy percentage. I ran a held-out evaluation, then used its failures to fix the extractor, which burns that holdout. Any number I quoted today would be inflated. A fresh untouched holdout is the next task. Until then every record carries a confidence tier and only the CONFIRMED tier is presented as fact.

– No spend or impression estimates. This answers who advertises where, not how much they paid.

Free tier is 200 requests/month, paid is $49/$199/$499. Keys are not self-serve yet, so the page is an early-access list rather than a checkout.

Two things I would actually like this sub’s read on: is a per-record evidence string useful to you, or is it dead weight next to a confidence score? And what would you want joined onto this that is missing (show category, audience estimates, historical backfill)?

https://podintel.github.io/?src=datasets

submitted by /u/Smart-Farmer1966
[link] [comments]

[PAID] I Built A UK Data API Platform With 37 Endpoints. One Token Balance, Every Dataset

Disclosure: I’m the developer and founder of StaticCreation. This is my own product.

Hey. Been working on this for over a year and just went live.

What it is: StaticCreation is an API marketplace for UK public sector data. Instead of scraping Companies House, DVSA, Land Registry, EPC, and a dozen other sources separately, you get one API key that works across everything.

The problem I kept hitting: Every time I built something that needed UK data, I’d spend weeks writing scrapers, handling rate limits, parsing inconsistent formats, and maintaining pipelines that break every time a government site changes. I figured other devs were doing the same thing.

What’s in there:

1.2 billion records across 16 datasets

Property: EPC certificates, Land Registry transactions, planning applications

Vehicles: DVLA data, full MOT history (832m test records), MIB insurance checks

Business: Companies House profiles, officers, charges, filings

Plus energy data, school/Ofsted ratings, crime stats, food data, fuel prices, trademarks, and more

How it works:

Buy tokens once (starts at £1.50 for 100), spend them per call. No subscriptions, no separate plans per dataset. One balance works across every endpoint.

Combined endpoints are what I’m most proud of, instead of calling 5 APIs yourself and joining the data, one call to Property Intelligence returns EPC + crime + schools + energy + price history for any UK address. Same idea for vehicles, companies, etc.

Stack: FastAPI backend, PostgreSQL, self-hosted on dedicated hardware in the UK. Everything runs through my own ETL pipelines, no reselling third-party APIs.

What I’d love feedback on:

Is the pricing clear enough?

Are there UK datasets you’d want that I’m missing?

Would you actually use something like this?

Site: https://staticcreation.co.uk

Happy to answer any questions about the build, the data partnerships, or the tech.

submitted by /u/Puzzleheaded_Bad_562
[link] [comments]

Watermarking Data Assets (Samples And Files)

QUESTION.

Is there a good way to watermark data assets before sharing with potential buyers?

We regularly share data samples with customers for evaluation, with clear licence terms on usage scope. But I worry those terms are practically unenforceable. Someone could generate synthetic data from a sample even though the licence restricts use to evaluation only.

Has anyone found effective ways to tag or watermark files before sharing? Metadata tagging is one option, but are there any deeper level solutions (steganographic watermarking, fingerprinting, etc)?

To keep it simple, let’s say we only talking about CSV files.
But this applies to video, audio, PDF, and archives too if you got any experience.

submitted by /u/Winter-Lake-589
[link] [comments]

[self-promotion] Read The Places: 2,105 Geocoded Real-world Places From 392 Novels, With Per-place Certainty Ratings And Source Passages (CC BY-SA 4.0)

Disclosure: this is my own project — I built and maintain it.

Source (the data itself): https://github.com/markselby9/readtheplaces.com — one directory per book under /books, each containing book.json, waypoints.json and source.txt.

Browsable version: https://readtheplaces.com

Scale: 392 novels, 2,105 places, 298 cities.

Per-place schema (waypoints.json), one real record, abridged:

{ "id": "westminster-doorstep", "name": "Clarissa's house, Westminster", "progressLabel": "10:00", "character": "clarissa", "coords": [-0.1275, 51.4993], "placeCertainty": "inferred", "certaintyNote": "Woolf never gives an address. The Dalloways live in Westminster within earshot of Big Ben; scholars place the house around Dean's Yard. Sited here as a considered guess, not a fact.", "quoteAnchor": "Mrs. Dalloway said she would buy the flowers herself.", "passage": "...", "sources": [...] } 

The field worth arguing about is placeCertainty. Geocoding fiction is mostly a disambiguation problem: many places are described rather than named (the abbey in The Name of the Rose is a northern Italian abbey Eco never names), and the named ones collide constantly. So each record carries what the resolution was based on, and inferred sitings say so in plain English instead of sitting on the map looking like facts. Filter to placeCertainty != "inferred" and you get a much smaller, much harder subset.

Waypoints are ordered by narrative progression rather than geography, so it’s usable for route/sequence work as well as point work.

How it was built, honestly: candidate mentions are extracted from the text by an LLM pass, then resolved against gazetteer data and checked by hand. Recall on minor mentions is therefore better than precision, and coverage skews heavily to 19th–20th century English-language fiction. Treat it as a curated dataset with a machine-assisted first pass, not a gold standard. It is not synthetic — every record points at a real passage in a real book.

Licence is CC BY-SA 4.0. Corrections are PRs against the JSON files, or there’s an issue template if you’d rather just report one.

submitted by /u/markselby9
[link] [comments]