Category: Datatards

Here you can observe the biggest nerds in the world in their natural habitat, longing for data sets. Not that it isn’t interesting, i’m interested. Maybe they know where the chix are. But what do they need it for? World domination?

When To Worry About Data Contamination In LLM Experiments?

Hey, I am currently preparing my master thesis experiment and was looking for datasets. My experiment will use LLMs as baseline with different RAG variations. Data contamination is a big topic for LLMs, because if the LLM has already been trained on the data I want use, then the whole experiment is pointless. The dataset I found on zenodo.org is for vulnerability detection.

Public and readable datasets are problematic, but what’s about downloadable datasets that do not have a preview on its side?

Should I be worried ?

submitted by /u/Apprehensive_Win662
[link] [comments]

Support Requested – RavenPack & Competitor Dataset Information

Hi all,

I’m helping a client evaluate a list of various data providers, but can’t quite seem to get a demo with some of these companies. It’s likely because their qualification process vets me out.

Is anyone willing to share the pricing of RavenPack’s products (like their sentiment analysis) the quality of their data?

If you have experience with other data providers, would love to learn about your experience with them as well.

Thanks in advance!

submitted by /u/AriCatalyx
[link] [comments]

URGENT: In Need Of Eco-home Dataset For My Thesis

Hi everyone,

I am in need of a dataset that shows data about eco-home owners**.**

Definition of eco-home: “An eco-home, or eco-house, is a home built to have a minimal impact on the environment. Eco-homes are designed to use less energy, reduce waste, and conserve water.”

Does anyone know where I could potentially find something like this? I am at a dead end and have procrastinating badly due to being unable to find a data source for this niche.

Thanks so much.

*Edit: data around LEED-certified homes would work too

submitted by /u/DrivenCleats
[link] [comments]

Looking For Genome Data For A Hobby Project

So I am reading a lot about evolution and for a big part, that’s about genes. I’m now a few books down, so I can kind of confidently talk about those subjects now, but the thing is that I have never ever worked with or even explored genetic data. Mind you, I am a data scientist. As a hobby project, I want to explore some genetic datasets. Does anyone know of any good a freely available resources, or could someone tell me a little about the different types of genetic data?

submitted by /u/Necessary-Oil-353
[link] [comments]

Dataset Of 180-degree Stereoscopic VR Videos For VR Video Upscaling And Synthesis.

Hi! I’ve done quite a bit of research trying to find datasets that fit the description above. Essentially, I’m working on an AI that can upscale 180-degree VR videos, preferably they’d be SBS. As a bit of a side project, I’d also like to work on an AI that has only one eye’s view as an input, and the other as an output. Essentially turning a 2D video into a 3D SBS video. Any help/leads would be appreciated. Thank you!

submitted by /u/starblasters8
[link] [comments]

Need Secondary Sources On Independent Contracting Vs. Employment Data And Advice On Collecting Primary Source Data

So, I’m trying to do research on whether one should be an independent contractor or an employee. This includes benefits, pay, work/life balance and a bunch of other stats. Do you know of any good secondary sources that can help me research this and do you have any advice on how to make my own survey (the survey doesn’t have to be on reddit)?

Also, if you know a good sub to ask this in, go ahead and point that out.

submitted by /u/GB819
[link] [comments]

Missing Airport Data For A Travel Project

Iโ€™m working on building a comprehensive travel spreadsheet and I have a section that contains a lot of airport data. Iโ€™m currently trying to find a comprehensive list of annual passenger traffic and if the airport is a domestic, regional, international, etc. I Ideally want to be able to pull data from IATA directly, but I canโ€™t seem to find a good way to do that. Iโ€™ve been searching through GitHub and I havenโ€™t found a dataset that contains this information yet. I am open to adding more info to the spreadsheet, so if you have any other good data sources to check out regarding airports that would be great too!

submitted by /u/725525
[link] [comments]

Looking For News API For At Least The Last 20 Years

Hey all,

I hope this is the right forum, but I am kind of new to all of this.

I am looking for a news API (doesn’t really matter which type of API) which goes back to at least 2000.
Can be from one big (NYT or so source), but the more sources it covers the better. Must include financial news (but doesnt have to be limited to that)
Doesn’t have to be free (sure, the less the better)

I found a couple, but none of them goes further than let’s say the past 5 years.

Any help?

Cheers ๐Ÿ™‚

Edit: with financial news I don’t necessarily mean it very specific. Let’s say the API just Covers different newspaper, which have a financial section, that would be enough

submitted by /u/dsdxb
[link] [comments]

Dataset Copyright From Webscraping Issues

If I webscraped data from a website that ‘surveys’ users to populate their database, then publicly displays it for users to see without any paywall or sign up required, can I freely post and use this data as I please? I would like to make it publicly available, but I don’t want to infringe on anything while doing so.

My end goal would be to just post it on kaggle for public use as well as do some analysis viewable in some sort of website or dashboard

submitted by /u/megemann
[link] [comments]

What Stats For Analysing Healthcare Large Datasets For Prison And Mental Health

Hi everyone,

Hope youโ€™re all well, Iโ€™m in the early stages of designing a PhD project and hope to work with linked large datasets to evaluate mental healthcare in prison and forensic settings, and evaluate economic aspects and effectiveness of care. Iโ€™m hoping to base this work on linked datasets. So far Iโ€™ve been reading about the solutions for missing data, and been surprised at the number of theories. Really interesting stuff!

If anyone has any suggestions for how to approach this topic, or ideas for methods , resources, books, YouTube and general thoughts please these would all be really appreciated. Iโ€™m literally starting from scratch with the stats knowledge so grateful for any suggestions,

I see this as part of the background work rather than requesting anything unscrupulous!

Thank you in advance

submitted by /u/Ok_Plant8421
[link] [comments]

Resume/CV Dataset For A Smart-Recruiter Project

I’m looking for a large resume/CV dataset for my Smart-Recruiter project. I’m unable to find a suitable one on neither of the popular platforms like Kaggle or Google Dataset Search or UCI Machine Learning Repo.

Requirements:

Simple 1/2 pages of files. Preferred file type is PDF but anything will work right now. Trying to avoid dummy data.

P.S.: I found a dataset on Kaggle that has about 228 docx files but the problem with this dataset is it’s too long, like each docx file contains at least 6 pages on average. And this is my understanding that any resume that is beyond 2 pages, don’t make it to the interview process.

I’m open to suggestions.

submitted by /u/SougatDey
[link] [comments]

Looking For A Recent Machine Learning Dataset, To Perform Regression, Classification.

Hello all, I’ve been tasked with finding a dataset for one of my courses. But can’t find any recent decent dataset to perform machine learning tasks. There’s also the constraint of having at least 50k samples and around 20 more or less features. I found some on kaggle but needed to delge more. Where can I look for more datasets where I can specify queries like these?

submitted by /u/CatSweaty4883
[link] [comments]

Facebook Friends Network Analysis: How To Gather Data

Hello! I am a humanities masters student with no coding background. I am trying to create a social network analysis of an individual Facebook page. Iโ€™ve found instructions from 2019-2021 on how to gather friend data using Selenium, but these tools no longer work. Iโ€™m getting quite frustrated trying to find solutions. At this point is the Facebook API at all conducive to this data gathering? Thank you in advance.

submitted by /u/Rhinestonecrowboy
[link] [comments]

Requesting Dataset For Drug-Drug Interaction Prediction

Hello ,
Iโ€™m currently working on a college research project on Drug-Drug Interaction Prediction using Knowledge Graph Embeddings and a Convolutional-LSTM Network. I came across the paper

– Drug-Drug Interaction Prediction Based on Knowledge Graph Embeddings and Convolutional-LSTM Network by *Md. Rezaul Karim, Michael Cochez, Joao Bosco Jares, Mamtaz Uddin, Oya Beyan, and Stefan Decker (Fraunhofer FIT, RWTH Aachen University, University of Dhaka).

If anyone has access to the dataset (or a similar one), or knows how I can obtain it, Iโ€™d really appreciate your help!

this would be really helpful .As i cant find the dataset from Kaggle also or from any source .

submitted by /u/__Silverfang__21
[link] [comments]

Help Creating A Deepfake Audio Dataset?

Hey everyone,

Iโ€™m working on building a deepfake audio dataset and wanted to get some help on best practices. I want to ensure that the dataset is diverse and representative for training an effective detection model.

Some questions I have:

How many speakers should I aim for to get a balanced dataset?

Should I maintain an equal gender ratio, or does it make a difference ?

How long is enough from each source(mins, hours)

Any recommended sources or strategies for collecting high-quality real audio?

What sample rates (e.g., 16kHz, 44.1kHz, 48kHz) or a what mix?

Are certain codecs (e.g., MP3, AAC, Opus, WAV) more challenging for detection models?

Would love to hear from those who have experience

submitted by /u/Fuzzy_Cream_5073
[link] [comments]

Open-MalSec V0.1 โ€“ Open-Source Cybersecurity / Analysis Samples

Evening! ๐Ÿซก

Just uploaded Open-MalSec v0.1, an early-stage open-source cybersecurity dataset focused on phishing, scams, and malware-related text samples.

๐Ÿ“‚ This is the base version (v0.1)โ€”just a few structured sample files. Full dataset builds will come over the next few weeks.

๐Ÿ”— Dataset link: huggingface.co/datasets/tegridydev/open-malsec

๐Ÿ” Whatโ€™s in v0.1?

A few structured scam examples (text-based)
Covers DeFi, crypto, phishing, and social engineering
Initial labelling format for scam classification

โš ๏ธ This is not a full dataset yet. Just establishing the structure + getting feedback.

๐Ÿ“‚ Current Schema & Labelling Approach

Each entry follows a structured JSON format with:

“instruction” โ†’ Task prompt (e.g., “Evaluate this message for scams”)
“input” โ†’ Source & message details (e.g., Telegram post, Tweet)
“output” โ†’ Scam classification & risk indicators

Sample Entry

json { “instruction”: “Analyze this tweet about a new dog-themed crypto token. Determine scam indicators if any.”, “input”: { “source”: “Twitter”, “handle”: “@DogLoverCrypto”, “tweet_content”: “DOGGIEINU just launched! Invest now for instant 500% gains. Dev is ex-Binance staff. #memecrypto #moonshot” }, “output”: { “classification”: “malicious”, “description”: “Tweet claims insider connections and extreme gains for a newly launched dog-themed token.”, “indicators”: [ “Overblown profit claims (500% ‘instant’)”, “False or unverifiable dev background”, “Hype-based marketing with no substance”, “No legitimate documentation or audit link” ] } }

๐Ÿ—‚๏ธ Current v0.1 Sample Categories

Crypto Scams โ†’ Meme token pump & dumps, fake DeFi projects

Phishing โ†’ Suspicious finance/social media messages

Social Engineering โ†’ Manipulative messages exploiting trust

๐Ÿ”œ Next Steps

๐Ÿ” Planned Updates:

Expanding dataset with more phishing & malware examples

Refining schema & annotation quality

Open to feedback, contributions, and suggestions

If this is useful, bookmark/follow the dataset here:

๐Ÿ”— huggingface.co/datasets/tegridydev/open-malsec

More updates coming as I expand the datasets ๐Ÿซก

๐Ÿ’ฌ Thoughts, feedback, and ideas are always welcome! Drop a comment or DMs are open ๐Ÿค™

submitted by /u/tegridyblues
[link] [comments]