submitted by /u/cavedave
[link] [comments]
Category: Datatards
Here you can observe the biggest nerds in the world in their natural habitat, longing for data sets. Not that it isn’t interesting, i’m interested. Maybe they know where the chix are. But what do they need it for? World domination?
Hi everyone,
Does anyone know of any panel datasets on video gaming (inc. mobile gaming) in the UK, or even just England and Wales? Looking to write a paper on video gaming habits.
I’ve seen reports on Statista, but you have to pay a hefty sum!
submitted by /u/Flimsy-sam
[link] [comments]
Where do I find previous years sales dataset for forecast
submitted by /u/Working-Tie-240
[link] [comments]
I’m looking for a large resume/CV dataset for my Smart-Recruiter project. I’m unable to find a suitable one on neither of the popular platforms like Kaggle or Google Dataset Search or UCI Machine Learning Repo.
Requirements:
Simple 1/2 pages of files. Preferred file type is PDF but anything will work right now. Trying to avoid dummy data.
P.S.: I found a dataset on Kaggle that has about 228 docx files but the problem with this dataset is it’s too long, like each docx file contains at least 6 pages on average. And this is my understanding that any resume that is beyond 2 pages, don’t make it to the interview process.
I’m open to suggestions.
submitted by /u/SougatDey
[link] [comments]
Hello all, I’ve been tasked with finding a dataset for one of my courses. But can’t find any recent decent dataset to perform machine learning tasks. There’s also the constraint of having at least 50k samples and around 20 more or less features. I found some on kaggle but needed to delge more. Where can I look for more datasets where I can specify queries like these?
submitted by /u/CatSweaty4883
[link] [comments]
I’m looking for a dataset that provides projections about the labor market by zip code. Ideally it would be for year 2023, but something as old as 2020 could suffice. I know the BLS only separates by state and I’m not seeing anything newer than 2018 from the US Census (doesn’t mean I’m not missing something).
Any help is appreciated!
submitted by /u/teacherofderp
[link] [comments]
Hey there, im looking for volleyball and rugby dataset. Is there any website with updated matches?
submitted by /u/Zealousideal-Key9042
[link] [comments]
Hello! I am a humanities masters student with no coding background. I am trying to create a social network analysis of an individual Facebook page. I’ve found instructions from 2019-2021 on how to gather friend data using Selenium, but these tools no longer work. I’m getting quite frustrated trying to find solutions. At this point is the Facebook API at all conducive to this data gathering? Thank you in advance.
submitted by /u/Rhinestonecrowboy
[link] [comments]
Hello ,
I’m currently working on a college research project on Drug-Drug Interaction Prediction using Knowledge Graph Embeddings and a Convolutional-LSTM Network. I came across the paper
– Drug-Drug Interaction Prediction Based on Knowledge Graph Embeddings and Convolutional-LSTM Network by *Md. Rezaul Karim, Michael Cochez, Joao Bosco Jares, Mamtaz Uddin, Oya Beyan, and Stefan Decker (Fraunhofer FIT, RWTH Aachen University, University of Dhaka).
If anyone has access to the dataset (or a similar one), or knows how I can obtain it, I’d really appreciate your help!
this would be really helpful .As i cant find the dataset from Kaggle also or from any source .
submitted by /u/__Silverfang__21
[link] [comments]
Any idea where to find raw email datasets?
I looked at the Eron dataset, but it is already cleaned I would like to improve my skills in data cleaning.
Any help would be greatly appreciated.
submitted by /u/mr_house7
[link] [comments]
Hey everyone,
I’m working on building a deepfake audio dataset and wanted to get some help on best practices. I want to ensure that the dataset is diverse and representative for training an effective detection model.
Some questions I have:
How many speakers should I aim for to get a balanced dataset?
Should I maintain an equal gender ratio, or does it make a difference ?
How long is enough from each source(mins, hours)
Any recommended sources or strategies for collecting high-quality real audio?
What sample rates (e.g., 16kHz, 44.1kHz, 48kHz) or a what mix?
Are certain codecs (e.g., MP3, AAC, Opus, WAV) more challenging for detection models?
Would love to hear from those who have experience
submitted by /u/Fuzzy_Cream_5073
[link] [comments]
I am working on a data analysis project but I’m having a difficult time find any datasets for Walmart Product Reviews with maybe 2022 or 2023 data. Any ideas?
submitted by /u/Zealousideal-Grab216
[link] [comments]
Evening! 🫡
Just uploaded Open-MalSec v0.1, an early-stage open-source cybersecurity dataset focused on phishing, scams, and malware-related text samples.
📂 This is the base version (v0.1)—just a few structured sample files. Full dataset builds will come over the next few weeks.
🔗 Dataset link: huggingface.co/datasets/tegridydev/open-malsec
🔍 What’s in v0.1?
A few structured scam examples (text-based)
Covers DeFi, crypto, phishing, and social engineering
Initial labelling format for scam classification
⚠️ This is not a full dataset yet. Just establishing the structure + getting feedback.
📂 Current Schema & Labelling Approach
Each entry follows a structured JSON format with:
“instruction” → Task prompt (e.g., “Evaluate this message for scams”)
“input” → Source & message details (e.g., Telegram post, Tweet)
“output” → Scam classification & risk indicators
Sample Entry
json { “instruction”: “Analyze this tweet about a new dog-themed crypto token. Determine scam indicators if any.”, “input”: { “source”: “Twitter”, “handle”: “@DogLoverCrypto”, “tweet_content”: “DOGGIEINU just launched! Invest now for instant 500% gains. Dev is ex-Binance staff. #memecrypto #moonshot” }, “output”: { “classification”: “malicious”, “description”: “Tweet claims insider connections and extreme gains for a newly launched dog-themed token.”, “indicators”: [ “Overblown profit claims (500% ‘instant’)”, “False or unverifiable dev background”, “Hype-based marketing with no substance”, “No legitimate documentation or audit link” ] } }
🗂️ Current v0.1 Sample Categories
Crypto Scams → Meme token pump & dumps, fake DeFi projects
Phishing → Suspicious finance/social media messages
Social Engineering → Manipulative messages exploiting trust
🔜 Next Steps
🔍 Planned Updates:
Expanding dataset with more phishing & malware examples
Refining schema & annotation quality
Open to feedback, contributions, and suggestions
If this is useful, bookmark/follow the dataset here:
🔗 huggingface.co/datasets/tegridydev/open-malsec
More updates coming as I expand the datasets 🫡
💬 Thoughts, feedback, and ideas are always welcome! Drop a comment or DMs are open 🤙
submitted by /u/tegridyblues
[link] [comments]
I`m trying to make a project with creating an OCR model for Ukrainian cursive recognition. I found one dataset with seperate Ukrainian letters, but I can`t fing a dataset with words, sentences, texts e.t.c. Help me please^(
submitted by /u/Klutzy-Translator-23
[link] [comments]
The dataset was processed and published on the Metabase BI platform.
It can be useful for research purposes.
Unfortunately, it’s closed under the simple registration as it might go down due to high load.
UK Dataset
submitted by /u/rzykov
[link] [comments]
What platforms can you get datasets from?
Instead of Kaggle and Roboflow
submitted by /u/Yennefer_207
[link] [comments]
[Sorry for my bad English. English is not my native language.]
Hello,
I am currently a student studying computer engineering. I need to do a graduation project in order to graduate. Since I have worked on NLP a lot before, I want my graduation project to be about NLP. I plan to develop a model that tries to identify the psychological disorders these people have, based on the writings written by people with psychological disorders.
However, I am having difficulty at the first stage. I have not been able to find a dataset to classify for a week. This is the only data set that can be useful to me, but it is not enough for me. reddit mental health data
I tried creating artificial datasets, but they didn’t give the results I wanted. What can I do about this?
Thank you very much in advance for your help.
submitted by /u/BaranKanat
[link] [comments]
I downloaded the 449M zip file that contains csv files. The branded_food.csv file has a column for the brand name but it’s bank. For example there are rows of products for PEPPERIDGE FARM but it’s not telling what products for PEPPERIDGE FARM.
Are there other sources I can download from which have more complete data?
I am looking for data like the nutritional label that’s in the back of every packaged food.
submitted by /u/THenrich
[link] [comments]
Looking for a large dataset of different foods spectral data, one containing nutritional information would be good, but the more datasets the better as I can just guess the rough nutritional data through the food name, as this isn’t for a precise purpose
submitted by /u/IsaacModdingPlzHelp
[link] [comments]
Just getting into data analytics and decided that I wanted to create my own project to practice. Looking for Portland, Oregon job market data. Hopefully something in the range of 2020 – 2024. Any suggestions or links?
submitted by /u/rconklin08
[link] [comments]
I need to replicate the below paper in which the dataset in title has been used.
The paper: Goldwater, S., Jurafsky, D., & Manning, C. D. (2010). Which words are hard to recognize? Prosodic, lexical, and disfluency factors that increase speech recognition error rates. Speech Communication, 52(3), 181–200.
submitted by /u/fiveMop
[link] [comments]
Hi everyone,
I’m working on a research project comparing LLM-generated text with human-written text. Does anyone know of a validated dataset (with DOI) that includes both? If not, could you share tips on creating one?
LLM text: Best models/prompts to generate diverse samples? Human text: Reliable sources for high-quality text? Validation: How to ensure balance and avoid bias?
Any help or pointers would be greatly appreciated! Thanks in advance.
submitted by /u/National_Evidence548
[link] [comments]
Hello, I want to make a website using Trader Joe’s products. Is there any way to access the list directly through their website? Otherwise, are there any public datasets? I just need information like the product name and picture.
submitted by /u/pcanpie
[link] [comments]
Hi , I’m currently working on a Food Nutrition App for my final year project , I’m having a hard time finding datasets of food with their nutritional values including pictures . Please help if you have any suggestions for website .
submitted by /u/No_Archer_9853
[link] [comments]
Like title, hoping for a recent dataset with a large amount of games, ideally from the premiere league. I wish for there to be player locations with each action, such as their location when they took a shot. Ideally it would be consistently updated, however that is not necessary.
For example I am looking for a dataset similar to the one used in this analysis:
https://www.kaggle.com/code/usamawaheed/expected-goals-xg-model/notebook
Thank you all
submitted by /u/BDubs5764
[link] [comments]
Please recommend free Historic Weather Datasets
submitted by /u/AcademicGuide997398
[link] [comments]
Where does this data come from?
Amazon.com features a best-sellers listing page for every category, subcategory, and further subdivisions.
I accessed each one of them. Got a total of 25,874 best seller pages.
For each page, I extracted data from the #1 product detail page – Name, Description, Price, Images and more. Everything that you can actually parse from the HTML.
There’s a lot of insights that you can get from the data. My plan is to make it public so everyone can benefit from it.
I’ll be running this process again every week or so. The goal is to always have updated data for you to rely on.
Where does this data come from?
Rating: Most of the top #1 products have a rating of around 4.5 stars. But that’s not always true – a few of them have less than 2 stars.
Top Brands: Amazon Basics dominates the best sellers listing pages. Whether this is synthetic or not, it’s interesting to see how far other brands are from it.
Most Common Words in Product Names: The presence of “Pack” and “Set” as top words is really interesting. My view is that these keywords suggest value—like you’re getting more for your money.
Raw data:
You can access the raw data here: https://github.com/octaprice/ecommerce-product-dataset.
Let me know in the comments if you’d like to see data from other websites/categories and what you think about this data.
submitted by /u/LessBadger4273
[link] [comments]
Hi, I’m looking for a dataset in CSV form that contains sequential game logs of player actions, either individual actions or completed goals (such as completing a level then moving on to the next level, quitting the game or choosing another activity within the game). I’m looking to build a model that predicts the action a player will take based on past in-game actions.
submitted by /u/RazorBeamer
[link] [comments]
I’m a bit confused about something with the [RAVDESS Emotional Speech Audio] dataset. I noticed that the file numbers on Kaggle don’t match the original dataset on Zenodo. From the original source, there should be 192 files per class (spread across 8 emotions: Neutral, Calm, Happy, Sad, Angry, Fearful, Disgust, Surprised).
But in the Kaggle version:
Most classes (like Happy, Sad, etc.) have 384 files instead of 192.
Two classes (Neutral and Calm) have around 2544 files, which is a lot more than expected.
Has anyone else noticed this? Could this be due to changes made by the uploader, or is there another reason? Would love to hear if anyone has more context!
submitted by /u/lama_777a
[link] [comments]