Category: Datatards

Here you can observe the biggest nerds in the world in their natural habitat, longing for data sets. Not that it isn’t interesting, i’m interested. Maybe they know where the chix are. But what do they need it for? World domination?

Start Golf Season With 90 Days Of Free PGA API Access (Free Giveaway)

Hey Reddit! 👋

With the PGA season heating up, we’re giving away 90 days of free access to our PGA API to the first 20 people who sign up by Sunday, February 9th. This isn’t a sales pitch—there’s no commitment, no credit card required—just an opportunity for those of you who love building, experimenting, and exploring with sports data.

Here’s what you’ll get access to:

Real-time tournament stats Past tournament stats Season schedules, golfer information + more

Curious about the API? You can check out the full documentation here: PGA API Documentation

We know there are tons of creative developers, analysts, and data enthusiasts here on Reddit who can do amazing things with access to this kind of data, and we’d love to see what you come up with. Whether you’re building an app, testing a project, or just curious to explore, this is for you.

If you’re interested, join our discord to sign up – just let us know you’re joining for PGA data! Spots are limited to the first 20, so don’t wait too long!

We’re really excited to see how you’ll use this. If you have any questions, feel free to ask in the comments or DM us.

submitted by /u/rollinginsights
[link] [comments]

VGGSound – Impossbile To Download Videos

Hi,

Navigating the complexities of dataset acquisition for my PhD research has proven challenging, particularly with the VGGSound dataset. Despite my extensive efforts, I’ve encountered significant roadblocks in downloading the required audio files. While the GitHub repository speedyseal/audiosetdl suggests a straightforward download method with the command python download_audioset.py, both for VGGSound and audioSet, the actual video retrieval has been thwarted by unavailable resources. Ironically, recent ICLR 2024 publications reference this dataset.

If anyone can help, that would be awesome. Thanks

submitted by /u/Motor-Bobcat-3555
[link] [comments]

Image Dataset Benchmarking – Request For Comment

Hey there! We’re working on annotating a significant dataset of approximately 180M photography images complete with Exif and geolocation data and are exploring popular benchmarks in order to showcase the datasets value. What benchmarks would be helpful for the community in terms of showing the relative value of the dataset vs others? If you’re interested, here’s a sample of the dataset.

submitted by /u/EmetResearch
[link] [comments]

High School AP Research Project: Need Help Replacing Pushshift API For Reddit Data Collection

Hi everyone,

I’m a high school student working on my AP Research project, and I’m running into some issues with data collection that I could really use help with. My study focuses on analyzing how Reddit-driven stock recommendations impact long-term investment decisions. I’m specifically looking at subreddits like r/wallstreetbets, r/stock, r/investing, and r/SecurityAnalysis to track sentiment around different stocks and see if that sentiment can predict stock performance over time.

I had originally planned to use the Pushshift API to collect historical Reddit data, but with Reddit’s recent API changes, Pushshift no longer works. Since I’m pretty new to programming and APIs, I’m not sure what the best alternative is. I’ve tried looking into PRAW, but I’m concerned about its limitations when it comes to accessing older posts.

Here’s what I need:

A reliable way to collect historical Reddit posts (from 2022 to 2025 if possible). Advice on whether PRAW can handle this, or if there’s another tool or method I should use. Suggestions for workarounds or public datasets that might help with historical Reddit data.

Since this is part of a project I hope to eventually publish, I’m really eager to find a solution. I’d love any advice, resources, or guidance you can offer, especially considering I’m new to this and learning as I go.

Here’s a link to my original methodology plan if it helps clear up some questions. Feel free to add coments to the document!

Methodology Plan

submitted by /u/Immediate-Today-8157
[link] [comments]

[Synthetic] Synthetic Emotions: AI-Generated Videos Of Human Expressions

I am excited to share Synthetic Emotions, a dataset featuring AI-generated videos of individuals expressing different emotions, including happiness, anger, sadness, fear, surprise, disgust, love, confusion, and more.

This dataset was created using OpenAI Sora and consists of 100 short videos, each 5 seconds long, 480p resolution, 9:16 aspect ratio, and generated in one-shot to ensure consistency. The dataset covers a diverse range of ethnicities and demographics to provide a balanced representation of human emotions.

Key Details:

Video Duration: 5 seconds Resolution: 480p Aspect Ratio: 9:16 Generation Mode: One-shot using OpenAI Sora Total Videos: 100 Emotion Categories (10 total): Happiness and Joy, Anger, Sadness, Fear, Surprise, Disgust, Love and Affection, Confusion, Neutral/Everyday, Mixed Emotions

Potential Applications:

Emotion Recognition Research Affective Computing & AI-Human Interaction Synthetic Video Data Exploration

If you are working in emotion recognition, AI-human interaction, or affective computing, or are simply interested in how AI-generated human emotions compare to real-world expressions, this dataset may be useful.

The dataset is available on Hugging Face:
🔗 https://huggingface.co/datasets/aadityaubhat/synthetic-emotions

submitted by /u/aadityaubhat
[link] [comments]

When To Worry About Data Contamination In LLM Experiments?

Hey, I am currently preparing my master thesis experiment and was looking for datasets. My experiment will use LLMs as baseline with different RAG variations. Data contamination is a big topic for LLMs, because if the LLM has already been trained on the data I want use, then the whole experiment is pointless. The dataset I found on zenodo.org is for vulnerability detection.

Public and readable datasets are problematic, but what’s about downloadable datasets that do not have a preview on its side?

Should I be worried ?

submitted by /u/Apprehensive_Win662
[link] [comments]

Support Requested – RavenPack & Competitor Dataset Information

Hi all,

I’m helping a client evaluate a list of various data providers, but can’t quite seem to get a demo with some of these companies. It’s likely because their qualification process vets me out.

Is anyone willing to share the pricing of RavenPack’s products (like their sentiment analysis) the quality of their data?

If you have experience with other data providers, would love to learn about your experience with them as well.

Thanks in advance!

submitted by /u/AriCatalyx
[link] [comments]

URGENT: In Need Of Eco-home Dataset For My Thesis

Hi everyone,

I am in need of a dataset that shows data about eco-home owners**.**

Definition of eco-home: “An eco-home, or eco-house, is a home built to have a minimal impact on the environment. Eco-homes are designed to use less energy, reduce waste, and conserve water.”

Does anyone know where I could potentially find something like this? I am at a dead end and have procrastinating badly due to being unable to find a data source for this niche.

Thanks so much.

*Edit: data around LEED-certified homes would work too

submitted by /u/DrivenCleats
[link] [comments]

Looking For Genome Data For A Hobby Project

So I am reading a lot about evolution and for a big part, that’s about genes. I’m now a few books down, so I can kind of confidently talk about those subjects now, but the thing is that I have never ever worked with or even explored genetic data. Mind you, I am a data scientist. As a hobby project, I want to explore some genetic datasets. Does anyone know of any good a freely available resources, or could someone tell me a little about the different types of genetic data?

submitted by /u/Necessary-Oil-353
[link] [comments]

Dataset Of 180-degree Stereoscopic VR Videos For VR Video Upscaling And Synthesis.

Hi! I’ve done quite a bit of research trying to find datasets that fit the description above. Essentially, I’m working on an AI that can upscale 180-degree VR videos, preferably they’d be SBS. As a bit of a side project, I’d also like to work on an AI that has only one eye’s view as an input, and the other as an output. Essentially turning a 2D video into a 3D SBS video. Any help/leads would be appreciated. Thank you!

submitted by /u/starblasters8
[link] [comments]

Need Secondary Sources On Independent Contracting Vs. Employment Data And Advice On Collecting Primary Source Data

So, I’m trying to do research on whether one should be an independent contractor or an employee. This includes benefits, pay, work/life balance and a bunch of other stats. Do you know of any good secondary sources that can help me research this and do you have any advice on how to make my own survey (the survey doesn’t have to be on reddit)?

Also, if you know a good sub to ask this in, go ahead and point that out.

submitted by /u/GB819
[link] [comments]

Missing Airport Data For A Travel Project

I’m working on building a comprehensive travel spreadsheet and I have a section that contains a lot of airport data. I’m currently trying to find a comprehensive list of annual passenger traffic and if the airport is a domestic, regional, international, etc. I Ideally want to be able to pull data from IATA directly, but I can’t seem to find a good way to do that. I’ve been searching through GitHub and I haven’t found a dataset that contains this information yet. I am open to adding more info to the spreadsheet, so if you have any other good data sources to check out regarding airports that would be great too!

submitted by /u/725525
[link] [comments]

Looking For News API For At Least The Last 20 Years

Hey all,

I hope this is the right forum, but I am kind of new to all of this.

I am looking for a news API (doesn’t really matter which type of API) which goes back to at least 2000.
Can be from one big (NYT or so source), but the more sources it covers the better. Must include financial news (but doesnt have to be limited to that)
Doesn’t have to be free (sure, the less the better)

I found a couple, but none of them goes further than let’s say the past 5 years.

Any help?

Cheers 🙂

Edit: with financial news I don’t necessarily mean it very specific. Let’s say the API just Covers different newspaper, which have a financial section, that would be enough

submitted by /u/dsdxb
[link] [comments]

Dataset Copyright From Webscraping Issues

If I webscraped data from a website that ‘surveys’ users to populate their database, then publicly displays it for users to see without any paywall or sign up required, can I freely post and use this data as I please? I would like to make it publicly available, but I don’t want to infringe on anything while doing so.

My end goal would be to just post it on kaggle for public use as well as do some analysis viewable in some sort of website or dashboard

submitted by /u/megemann
[link] [comments]

What Stats For Analysing Healthcare Large Datasets For Prison And Mental Health

Hi everyone,

Hope you’re all well, I’m in the early stages of designing a PhD project and hope to work with linked large datasets to evaluate mental healthcare in prison and forensic settings, and evaluate economic aspects and effectiveness of care. I’m hoping to base this work on linked datasets. So far I’ve been reading about the solutions for missing data, and been surprised at the number of theories. Really interesting stuff!

If anyone has any suggestions for how to approach this topic, or ideas for methods , resources, books, YouTube and general thoughts please these would all be really appreciated. I’m literally starting from scratch with the stats knowledge so grateful for any suggestions,

I see this as part of the background work rather than requesting anything unscrupulous!

Thank you in advance

submitted by /u/Ok_Plant8421
[link] [comments]