Looking For (US R1) Longitudinal Faculty Dataset

I’m looking for pointers to one or more datasets that have some or all of the following data:

Faculty name (tenure track only)
Current professional title/designation
Department employed
Name of the university/academic employer
Degree-granting department and institution (PhD, Masters, and undergraduate degrees, as applicable)
Year of degree (PhD, Masters, and undergraduate degrees)
Current employment start year
Other academic employment history (eg. department, start and end date of previous post-PhD employments)

It would be really nice if longitudinal data (every academic year) was also available for these items. In addition, data about non tenure track faculty appointments would also be nice, but not necessary.

I’m looking for something similar (but expanded in terms of scope) to the dataset used in this paper.

I’m aware that AARC could be a potential data source but I’ve been told it’s not trivial to get data access through them, so looking for alternatives.

Alternatively, would also appreciate if anyone can point me to ways to scrape (at least some of) this data from university directories.

Thanks in advance!

submitted by /u/Timely-Ad2743
[link] [comments]

0

Free [Synthetic] Datasets For AI Model Tuning [self-promotion]

I run a synthetic data platform called DataCreator AI that helps AI professionals and businesses generate customized datasets.

Along with these capabilities, we offer a section called Community Datasets where we post datasets for free. Community Datasets

Some of the current free datasets we have are:

A dataset to perform Direct Preference Optimization to reduce sycophancy of LLMs.
A dataset that contains structured multi-turn conversations between patients and customer service agents at hospitals.
A dataset with a collection of random facts from various topics like biology, astronomy,
Classification and Question-Answer Datasets.

Your feedback would be of huge help to me to come up with more useful datasets. If you have any specific dataset ideas, please let me know in the comments so that we can put up more of them.

submitted by /u/Routine-Sound8735
[link] [comments]

0

Can Someone Help Me Find The News Headlines Every Day For The Last 100 Days Please?

From the main worldwide news providers is great!

submitted by /u/Actual-Bid-853
[link] [comments]

0

Oral Health Buyers Demographics – Age

Hiya, I’m investigating marketing to oral health care companies and what to simply know how their market is segmented, by purchases, by age and sex.

General or specific info would be fine. I suspect it’s women, but what age range?

submitted by /u/RickNBacker4003
[link] [comments]

0

Help Needed: Collect 100–150 Samples Per Bird Species (Images + Audio) For Dataset

Hi everyone,
I’m working on a bird species classification + migration prediction project for my capstone. I have a list of ~512 bird species, and I need help collecting at least 100–150 samples per species (images, and audio if possible).

submitted by /u/Shrinivas-k-shreeni
[link] [comments]

0

Complete Powerball & Mega Millions Draw + Winners Dataset

I’m working on a data project and need a more complete dataset for Powerball and Mega Millions than what’s usually available on sites like lotteryusa or state lottery pages.

Most public datasets just have the draw date and winning numbers, but I need all the columns, specifically things like: – Draw date & draw number – Winning numbers + Powerball/Mega Ball – Power Play / Megaplier multiplier – Jackpot amount (annuity & cash value) – Number of winners by tier (match 5, 4+PB, etc.) – Power Play winners by tier – State-by-state winner breakdown (if available)

Basically, the full official results table that the lotteries publish after each draw, not just the numbers themselves.

I haven’t been able to find a historical dataset with all of this.

Does anyone know if this exists publicly, or will I need to scrape it directly from Powerball.com / MegaMillions.com (or individual state sites)? If scraping is the way to go, I’d love any tips on best practices for this since the data spans back to the ’90s.

submitted by /u/b2bdemand
[link] [comments]

0

(Urgent) Needd Advice For Dataset Creation

I have 90 videos downloaded from yt i want to crop them all just a particular section of the videos its at the same place for all the videos and i need its cropped video along with the subtitles is there any software or ml model through which i can do this quicklyy?

submitted by /u/courage10asd
[link] [comments]

0

Requesting Supply Chain Dataset For Academic Research

I am conducting academic research on supplier evaluation and selection using machine learning as part of my postgraduate work. For this, I am seeking access to supplier-related datasets that include features such as unit price, product availability, order quantities, revenue generated, stock levels, lead times, shipping times, shipping costs, shipping carriers, supplier location, production volumes, manufacturing lead times, manufacturing costs, defect rates, transportation modes, and overall procurement costs. The data will be used strictly for academic purposes, and any confidential or sensitive information will be anonymized. Access to such data would greatly enhance the reliability of my research and contribute to building a practical decision-support framework for procurement systems.
If these features are not there any dataset will do. Please I really need the dataset

submitted by /u/BackgroundFar8017
[link] [comments]

0

Survey For A Data Marketplace | For Anyone Looking To Earn From Data

I’m in the process of developing a marketplace to sell data because I feel like there is no simple marketplace to facilitate sell data, especially for subscriptions and I really wanted people in the communities opinions. If you have data, are interested in selling data etc. an entry would be appreciated, it has been checked by mods, emails are not collect

Here is the link: https://forms.gle/xNp7a7vEEioa7vrE8

submitted by /u/daviddosm8
[link] [comments]

0

Budget-friendly Alternatives For Grocery Product Datasets?

Looking for paid dataset providers for Indian grocery/retail data (similar to quick-commerce platforms).

Format: CSV/JSON

submitted by /u/Top_Sundae8258
[link] [comments]

0

New Analyst Building A Portfolio While Job Hunting-what Datasets Actually Show Real-world Skill?

I’m a new data analyst trying to land my first full-time role, and I’m building a portfolio and practicing for interviews as I apply. I’ve done the usual polished datasets (Titanic/clean Kaggle stuff), but I feel like they don’t reflect the messy, business-question-driven work I’d actually do on the job.

I’m looking for public datasets that let me tell an end-to-end story: define a question, model/clean in SQL, analyze in Python, and finish with a dashboard. Ideally something with seasonality, joins across sources, and a clear decision or KPI impact.

Datasets I’m considering: – NYC TLC trips + NOAA weather to explain demand, tipping, or surge patterns – US DOT On-Time Performance (BTS) to analyze delay drivers and build a simple ETA model – City 311 requests to prioritize service backlogs and forecast hotspots – Yelp Open Dataset to tie reviews to price range/location and detect “menu creep” or churn risk – CMS Hospital Compare (or Medicare samples) to compare quality metrics vs readmission rates

For presentation, is a repository containing a clear README (business question, data sources, and decisions), EDA/modeling notebooks, a SQL folder for transformations, and a deployed Tableau/Looker Studio link enough? Or do you prefer a short write-up per project with charts embedded and code linked at the end?

On the interview side, I’ve been rehearsing a crisp portfolio walkthrough with Beyz interview assistant, but I still need stronger datasets to build around. If you hire analysts, what makes you actually open a portfolio and keep reading?

Last thing, are certificates like DataCamp’s worth the time/money for someone without a formal DS degree, or would you rather see 2–3 focused, shippable projects that answer a business question? Any dataset recommendations or examples would be hugely appreciated.

submitted by /u/Various_Candidate325
[link] [comments]

0

Is It Possible To Make Decent Money Making Datasets With A Good IPhone Camera?

I can record videos or take photos of random things outside or around the house, label and add variations on labels. Where might I sell datasets and how big would they have to be to be worth selling?

submitted by /u/No-Yak4416
[link] [comments]

0

Guys I Need A Image Dataset Of Medical Forms

I need dataset of medical forms like medical reports, hospital admission form, medical insurance form,etc .

Please drop links

submitted by /u/Fit-Metal7779
[link] [comments]

0

Where To Find Good Relation Based Datasets?

Okay so I need to find a dataset that has at least like 3 tables, I’m search stuff on kaggle like supermarket or something and I can’t seem to find simple like a products table, order etc. Or maybe a bookstore I don’t know. Any suggestions?

submitted by /u/aphroditelady13V
[link] [comments]

0

Free Tool: Explore Facebook Ads Library Pages By Keywords And Other Filters

submitted by /u/firepost
[link] [comments]

0

Need Help In Predicting The Next Half Of A Dataset. There Will Be A Cash Reward For The First Person To Solve It

https://www.dropbox.com/scl/fi/vm7zztz460hfgb0sxy633/bounty-columns-offset-data-sample.csv?rlkey=ytsp9dcuabxhywhun5tbs1lm6&e=2&st=ogqkbbez&dl=0

this is the provided data set and i need someone to predict the next half of the dataset with either 90% or 100% accuracy please

I don’t care how you solve it, only that you provide proof of the solve, and the algo code that solved it. Must provide full code to replicate.

The data is multi-dimensional, and catalogued. I have both halves of the data, to compare against.

Thanks, dm me if you are interested, i am ready to offer upwards of 150 USD for the solution

submitted by /u/waduhek77
[link] [comments]

0

Where Can I Get Real-time Gas/fuel Price Data (API Or Dataset) In Canada?

Hi everyone,

I’m working on a side project and need real-time gas/fuel price data in Canada.

I know GasBuddy and Waze get theirs from crowdsourcing. GasBuddy also used to have a GraphQL API, but that seems shut down. I already emailed OPIS but got no response.

Ideally, I’m looking for:

Station-level data with location
Prices by fuel type (regular, premium, diesel, etc.)
Search by postal code or lat/long
Brand filtering if possible
Fuel price based on the type of fuel – Petrol, Diesel and also the price for Regular, Premium etc.

Are there any real-time APIs or datasets available for this? Or is scraping the only realistic option here for real-time data for the daily fuel price?

Thanks! 🙏

submitted by /u/Unhappy_Bug_5277
[link] [comments]

0

A Comprehensive List Of Open-source Datasets For Voice And Sound Computing (95+ Datasets).

submitted by /u/cavedave
[link] [comments]

0

The Worlds 2.7B Buildings Geodata From The Munich.

submitted by /u/cavedave
[link] [comments]

0

ML Data Pipeline Pain Points Whats Your Biggest Preparing Frustration?

Researching ML data pipeline pain points. For production ML builders: what’s your biggest training data prep frustration?

🔍 Data quality? ⏱️ Labeling bottlenecks? 💰 Annotation costs? ⚖️ Bias issues?

Share your real experiences!

submitted by /u/3DMakeorg
[link] [comments]

0

Anybody Else Running Into This Problem With Datasets?

Spent weeks trying to find realistic e-commerce data for AI/BI testing, but most datasets are outdated or privacy-risky. Ended up generating my own synthetic datasets — users, products, orders, reviews — and packaged them for testing/ML. Curious if others have faced this too?

https://youcancallmedustin.github.io/synthetic-ecommerce-dataset/

submitted by /u/ItsThinkBuild
[link] [comments]

0

📊 New Dataset: 2.6M+ AI-enriched Company Profiles Across 100+ Industries (JSONL / Parquet / CSV)

Hi all,

I’ve been working on a side project where I crawled and AI-enriched over 2.6 million company websites across 111 industries worldwide.

What’s inside:

Company name, website, industry
Long + short descriptions (AI-generated)
Enriched metadata (socials, emails, locations where available)
Website screenshots
Delivered in JSONL, Parquet, and CSV formats

Access:

A free sample explorer with 150 companies is live here: https://ctxdb.ai/sample-dataset
Full dataset available for purchase (Q3 2025 edition + Q4 coming soon).
A yearly “Momentum Plan” also refreshes the dataset quarterly with new companies + updated profiles.

Why I built this:

I wanted an up-to-date, structured dataset useful for:

Lead generation / prospecting
Market research & competitive tracking
AI/ML model training
Academic or investment research

Happy to hear your thoughts / feedback / need for API access? – also curious how you’d use a dataset like this.

submitted by /u/karngyan
[link] [comments]

0

New Mapping Created To Normalize 11,000+ XBRL Taxonomy Names For Better Financial Data Analysis

Hey everyone! I’ve been working on a project to make SEC financial data more accessible and wanted to share what I just implemented. https://nomas.fyi

**The Problem:**

XBRL taxonomy names are technical and hard to read or feed to models. For example:

– “EntityCommonStockSharesOutstanding”

These are accurate but not user-friendly for financial analysis.

**The Solution:**

We created a comprehensive mapping system that normalizes these to human-readable terms:

– “Common Stock, Shares Outstanding”

**What we accomplished:**

✅ Mapped 11,000+ XBRL taxonomies from SEC filings

✅ Maintained data integrity (still uses original taxonomy for API calls)

✅ Added metadata chips showing XBRL taxonomy, SEC labels, and descriptions

✅ Enhanced user experience without losing technical precision

**Technical details:**

– Backend API now returns taxonomy metadata with each data response

– Frontend displays clean chips with XBRL taxonomy, SEC label, and full descriptions

– Database stores both original taxonomy and normalized display names

– Caching system for performance

Upvote1Downvote0Go to comments

submitted by /u/ccnomas
[link] [comments]

0

What Is Data Authorization And How To Implement It

submitted by /u/West-Chard-1474
[link] [comments]

0

Where Can I Find Dataset For Autism.

Hello there !

I am trying to find dataset for autism detection using EEG.
Can anyone link any source or anything.

Thanks…

submitted by /u/Available-Fee1691
[link] [comments]

0

Suggestions And Recommendations For Creating A Custom Dataset For Fine Tuning A LLM

submitted by /u/Old-Raspberry-3266
[link] [comments]

0

I Built A Daily Startup Funding Dataset (updated Daily) – Feedback Appreciated!

Hey everyone!

As a side project, I started collecting and structuring data on recently funded startups (updated daily). It includes details like:

Company name, industry, description
Funding round, amount, date
Lead + participating investors
Founders, year founded, HQ location
Valuation (if disclosed) and previous rounds

Right now I’ve got it in a clean, google sheet, but I’m still figuring out the most useful way to make this available.

Would love feedback on:

Who do you think finds this most valuable? (Sales teams? VCs? Analysts?)
What would make it more useful: API access, dashboards, CRM integration?
Any “must-have” data fields I should be adding?

This started as a freelance project but I realized it could be a lot bigger, and I’d appreciate ideas from the community before I take the next step.

Link to dataset sample – https://docs.google.com/spreadsheets/d/1649CbUgiEnWq4RzodeEw41IbcEb0v7paqL1FcKGXCBI/edit?usp=sharing

submitted by /u/Capable_Atmosphere_7
[link] [comments]

0

Huge Open-Source Anime Dataset: 1.77M Users & 148M Ratings

Hey everyone, I’ve published a freshly-built anime ratings dataset that I’ve been working on. It covers 1.77M users, 20K+ anime titles, and over 148M user ratings, all from engaged users (minimum 5 ratings each).

This dataset is great for:

Building recommendation systems
Studying user behavior & engagement
Exploring genre-based analysis
Training hybrid deep learning models with metadata

🔗 Links:

Kaggle Dataset: https://www.kaggle.com/datasets/tavuksuzdurum/user-animelist-dataset (inference notebook available)
Hugging Face Space: https://huggingface.co/spaces/mramazan/AnimeRecBERT
GitHub Project (AnimeRecBERT Hybrid): https://github.com/MRamazan/AnimeRecBERT-Hybrid

submitted by /u/RealisticGround2442
[link] [comments]

0

[self-promotion] Free Sample: EU Public Procurement Notices (Aug 2025, CSV, Enriched With CPV Codes)

I’ve released a new dataset built from the EU’s Tenders Electronic Daily (TED) portal, which publishes official public procurement notices from across Europe.

Source: Official TED monthly XML package for August 2025
Processing: Parsed into a clean tabular CSV, normalized fields, and enriched with CPV 2008 labels (Common Procurement Vocabulary).
Contents (sample):
- notice_id — unique identifier
- publication_date — ISO 8601 format
- buyer_id — anonymized buyer reference
- cpv_code + cpv_label — procurement category (CPV 2008)
- lot_id, lot_name, lot_description
- award_value, currency
- source_file — original TED XML reference

This free sample contains 100 rows representative of the full dataset (~200k rows).
Sample dataset on Hugging Face

If you’re interested in the full month (200k+ notices), it’s available here:
Full dataset on Gumroad

Suggested uses: training NLP/ML models (NER, classification, forecasting), procurement market analysis, transparency research.

Feedback welcome — I’d love to hear how others might use this or what extra enrichments would be most useful.

submitted by /u/OpenMLDatasets
[link] [comments]

0

Combining Parquet For Metadata And Native Formats For Video, Audio, And Images With DataChain AI Data Warehouse

The article outlines several fundamental problems that arise when teams try to store raw media data (like video, audio, and images) inside Parquet files, and explains how DataChain addresses these issues for modern multimodal datasets – by using Parquet strictly for structured metadata while keeping heavy binary media in their native formats and referencing them externally for optimal performance: reddit.com/r/datachain/comments/1n7xsst/parquet_is_great_for_tables_terrible_for_video/

It shows how to use Datachain to fix these problems – to keep raw media in object storage, maintain metadata in Parquet, and link the two via references.

submitted by /u/thumbsdrivesmecrazy
[link] [comments]

0

Category: Datatards