Category: Datatards

Here you can observe the biggest nerds in the world in their natural habitat, longing for data sets. Not that it isn’t interesting, i’m interested. Maybe they know where the chix are. But what do they need it for? World domination?

Study Roadmap/orientation App: Requesting For Data Fill (or At Least Sources). You Can Help Though The UI

The app is for students that help them:

– find which study programs you can apply for based on your exam track

– or discover careers you’re interested in and the study programs related to them 🙂

## Data

The data is still incomplete due to the lack of clear sources, but you can get what’s already there (Madagascar datasets are the only ones for now).

But mostly, you’re welcome to contribute 🙂

Repo: https://github.com/gigasandwich/giga-roadmap

Data (json) are stored in `/data`

App: https://roadmap.gigas.app

submitted by /u/_giga_sss_
[link] [comments]

[request] Embase Search Export (RIS/CSV) Help Required Pleaseee

Hi everyone) I’m doing a systematic review and unfortunately don’t have institutional access to embase.
can some please run the search for my and export the results in csv/RIS format?

The search prompt is:

(“Boron Neutron Capture Therapy”[MeSH] OR “boron neutron capture therapy”[Title/Abstract] OR BNCT[Title/Abstract])

Thank you so much!

submitted by /u/333afos
[link] [comments]

[PAID] Unified API For NOAA/ECMWF Weather, Climate, Aviation And Air-quality Datasets

Disclosure: I built and run GribStream, so this is self-promotion. It is a commercial/freemium API, but there is a free tier, and new free accounts currently get an intro quota boost so people can run real tests before deciding if it is useful.

The data itself is not mine. The original sources are public weather and climate feeds from NOAA/NCEP, ECMWF, Copernicus/ERA5, NOMADS, and public cloud archives. GribStream is a unified API and indexing layer on top of those datasets.

The idea is simple: instead of learning a different access pattern for every weather model feed, you can query many of them through the same API.

It covers common, broad-use datasets like:

  • GFS global forecasts
  • IFS deterministic and ensemble forecasts
  • HRRR high-resolution US forecasts
  • NBM forecast blends
  • ERA5 reanalysis
  • RTMA/URMA surface analyses
  • GEFS ensemble forecasts

And also more specialized datasets, for example:

  • AQM/NAQFC air-quality guidance for ozone and PM2.5
  • GTGN aviation turbulence nowcasts
  • aviation icing and turbulence feeds
  • SPC severe-weather probability products
  • wave, chemistry, UV index, and seasonal forecast products
  • AI weather model feeds and archives like AIFS, AIGFS/AIGEFS, GraphCastGFS, and FourCastNetGFS

What the API is mainly useful for:

  • pulling time series for one point or thousands of points
  • comparing forecast model runs over time
  • backtesting with “what was known at the time” cutoffs
  • querying multiple variables, levels, model runs, and ensemble members
  • getting data back as JSON, CSV, or NDJSON without building a custom weather-data pipeline first

A practical example: if you have 500 solar sites, farms, airports, ships, stores, or insurance exposure locations, you can ask for historical forecasts or recent model data at those coordinates directly, instead of separately wrangling GFS, HRRR, NBM, IFS, ERA5, etc.

Links:

GribStream: https://gribstream.com/

Model catalog: https://gribstream.com/models

Original/public source examples:

https://nomads.ncep.noaa.gov/

https://registry.opendata.aws/collab/noaa/

https://www.ecmwf.int/en/forecasts/dataset/open-data

https://cds.climate.copernicus.eu/

I’d be grateful for feedback from people who use weather, climate, aviation, energy, agriculture, logistics, or environmental datasets. Are there public weather datasets you wish were easier to query? Would you be looking for bulk exports? Interested in being able to setup notifications for weather data events?

Happy to answer questions. I’m trying to make the public model data easier to use while still being clear about where the original data comes from.

submitted by /u/ElPeque222
[link] [comments]

Does Anyone Have A Mirror For The FICS-PCB Dataset (TrustHub)? Deep Learning Hardware Assurance Project Stalled By Data Starvation.

I am building a YOLO-based PCB reverse engineering pipeline for academic research. I desperately need the FICS-PCB dataset (hosted on TrustHub) to scale my component detection model, but it is locked behind an authentication wall. I have emailed the authors but am waiting on a response. Looking for a mirror or anyone with TrustHub access.

The Engineering Context

I am currently working on an automated Printed Circuit Board (PCB) reverse engineering and hardware assurance pipeline. The end goal is automated Bill of Materials (BoM) extraction.

Initially, I replicated the classical image processing pipeline from Kleber et al. (2017). While I got decent IC detection using rigid OpenCV heuristics (HSV masking, Otsu thresholding, morphological transformations), the pipeline was far too brittle. Any change in PCB substrate color or environmental lighting required manual parameter tuning.

I recently pivoted the component detection stage to a Convolutional Neural Network (Ultralytics YOLO).

The Problem: Data Starvation

The YOLO architecture completely bypassed the need for manual CV parameter tuning and successfully isolated primary SoCs (97%+ confidence) against complex backgrounds. However, I am hitting a massive data starvation wall.

To make this model generalize across edge cases and minority classes, I need high-volume, annotated data.

The Roadblock

The FICS-PCB: A Multi-Modal Image Dataset (Lu et al., 2020) is exactly what I need. It contains 9,912 PCB sample images and over 77,000 component annotations.

  • The dataset is hosted onTrustHub.
  • It is locked behind an ID/password authentication barrier.
  • I have already sent a formal request from my institutional email to the principal investigators (University of Florida), but I am waiting on approval and my research sprint is currently bottlenecked.

submitted by /u/Specialist_Heron_906
[link] [comments]

Image Metadata Dataset For EXIF/IPTC/XMP Analysis, Provenance Research, And Forensic Workflows

I’m sharing an interest in datasets related to image metadata — especially EXIF, IPTC, and XMP fields — for use in forensic analysis, provenance research, search/indexing, and large-scale metadata extraction workflows.

I’m specifically looking for datasets that include one or more of the following:

  • Original image files with metadata intact.
  • Paired image + metadata exports.
  • Large collections suitable for testing extraction, indexing, normalization, or deduplication pipelines.
  • Real-world examples that include camera data, timestamps, geotags, creator info, editing history, and embedded tags.
  • Datasets useful for studying metadata loss across platforms or image-processing tools.

If anyone knows of public datasets, archives, or research corpora in this area, I’d appreciate recommendations. I’m especially interested in datasets that are legal to analyze and can be used for technical experimentation.

Disclosure: I work on image-meta.com, which is relevant to this topic.

submitted by /u/cstadler
[link] [comments]

I Published Free Samle Uniswap V3 BTC/ETH Research Datasets On Kaggle: Raw Logs, Swaps, 1-minute Bars, Liquidity Events, And Daily State Snapshots

I recently published two free Ethereum Uniswap V3 BTC/ETH datasets on Kaggle for researchers, quants, data scientists, and anyone studying DEX market structure.

These are not just price CSVs. The datasets include multiple research layers built from Ethereum mainnet data:

  • raw Uniswap V3 logs
  • decoded / normalized swaps
  • canonical 1-minute OHLCV bars
  • Mint, Burn, Collect liquidity events
  • Flash events
  • pool initialization data
  • pool registry metadata
  • daily archive-state snapshots

The pool universe covers 24 Uniswap V3 BTC/ETH-related pools:

  • WBTC/USDC
  • WBTC/USDT
  • WBTC/WETH
  • WETH/USDC
  • WETH/USDT
  • WETH/DAI

Across the major fee tiers:

  • 0.01%
  • 0.05%
  • 0.30%
  • 1.00%

The 2021 Kaggle sample covers 2021-05-04 to 2021-12-31 and includes about:

  • 2.98M raw logs
  • 2.78M normalized swaps
  • 1.17M canonical 1-minute bars
  • 288K liquidity events
  • daily pool state snapshots

The June 2026 sample covers 2026-06-01 to 2026-06-30 and includes about:

  • 1.61M raw logs
  • 1.57M normalized swaps
  • 329K canonical 1-minute bars
  • 76K liquidity events
  • daily pool state snapshots

Possible research ideas:

  • BTC/ETH DEX microstructure
  • Uniswap V3 liquidity behavior
  • fee tier comparison
  • pool-level volume and spread behavior
  • swap flow and buy/sell imbalance
  • LP activity around volatility regimes
  • comparing 2021 Uniswap V3 launch-era behavior vs 2026 mature-market behavior

I also included starter notebooks so people can quickly inspect the Parquet files and start exploring without building a full Ethereum indexer.

The public Kaggle datasets are free samples. I also maintain a larger validated archive covering 2021-05-04 to 2026-06-30 with the same research layers. If any researchers, teams, funds, or data builders need the full historical range or custom extracts, feel free to reach out through Kaggle.

Hope this helps anyone working on DeFi data, market microstructure, or crypto time-series research.

2021 sample: https://www.kaggle.com/datasets/marvingozo/ethereum-uniswap-v3-btceth-2021-free-sample

June 2026 sample: https://www.kaggle.com/datasets/marvingozo/ethereum-uniswap-v3-btceth-june-2026

submitted by /u/Upset-Fly-454
[link] [comments]

Free International Historical Return Data Files

Getting good data is a big hurdle for retail investors. Reliable return histories are often locked behind thousand dollar a year subscriptions. But you can get a lot for free.

I put together a small return dataset covering developed-market stocks, sovereign bonds, interest rates, and currencies.

The goal is to consolidate the kinds of return series that are useful for testing global asset allocation strategies, especially those involving foreign equity, sovereign bonds, currency hedging, and excess returns.

The dataset includes 50+ years of coverage across several files. All available for free. Check it out!

https://github.com/birjusuketupatel/ReturnDataFiles/tree/main

Note: Reposting bc the mods removed post on original subreddit.

submitted by /u/NecessarySpread2592
[link] [comments]

Query To Get Clinical Dataset For ML Project

hi everyone

im trying to get access to the hirid dataset for a machine learning project but im stuck at the citi course requirement because my organization isnt listed in the available options

has anyone run into this before or knows how to proceed in this situation any help would be really appreciated

thanks in advance

submitted by /u/malfoy011
[link] [comments]

Looking For A Dataset Like BuiltWith

Hi, so I am searching for any freely available dataset that would have information on websites using email services from third-parties.
BuiltWith provides that, like which websites are actively using Brevo, Klaviyo or MailChimp etc, but they are too expensive.
Thanks.

submitted by /u/Fari1911
[link] [comments]

I Built A Free, Unified API For Searching 500,000+ Companies Across Lithuania, Latvia, And Estonia.

Hey everyone,

As a side project, I got really frustrated trying to navigate the different government business registries across the Baltics whenever I needed to check a company’s status, VAT number, or employee count.

To solve this, I downloaded all the raw open data from the Lithuanian, Latvian, and Estonian registries and built a unified, lightning-fast search engine and API on top of it.

Link: https://www.balticdata.eu

Right now, the API is completely free and open for developers to use. You can instantly search by name, registration code, or filter by active/liquidated status.

I’d love for you to try it out and let me know if it’s useful or if you find any bugs!

submitted by /u/Slavcik
[link] [comments]

NIH Exporter Downloads Constantly Time Out

I’ve been trying to download project information for a specific year from the NIH Exporter tool (https://reporter.nih.gov/exporter/projects) and any file larger than 10Mb just times out every time. I tried downloading from the browser, from a console using bash tools, nothing works. There is no scheduled eRA maintenance going on. Anyone knows of any tricks for this? Has anyone tried to download project data recently?

submitted by /u/el_cadorna
[link] [comments]