Category: Datatards

Here you can observe the biggest nerds in the world in their natural habitat, longing for data sets. Not that it isn’t interesting, i’m interested. Maybe they know where the chix are. But what do they need it for? World domination?

Random Object Detection Dataset For Machine Learning

So I am trying to train an AI to detect all the small miscellaneous stuff within a image, for example like keys,bottle cap, bottle, wrapping paper, broken glass, paper and I want to exclude larger items like chair, table, fan, sofa, etcs. This AI will first need to detect these items before picking them up via some mechanical system.

submitted by /u/GateCodeMark
[link] [comments]

Generate My Own Data For Fine-tuning. Thoughts/tips/feedback?

So much focus on better models, not nearly enough on better post training data. I recently came across Curator, open source tool for dataset generation and refinement. It seems promising for automating parts of the process, has anyone here tried it? Would love to hear thoughts!

Also curious—how do you all handle data generation? Any tools that have worked well please feel free to share

submitted by /u/Ambitious_Anybody855
[link] [comments]

Need Help Finding Data Research Project

I am in dire need of help finding a viable dataset for my research project. I am in my final semester of undergrad and have been tasked with a major research project which will soon need to be transferred into STATA but for now, I need to run basic descriptive statisitcs and come up with my hypothesis, research question, and equation. No matter what topic I bounce around I can’t seem to find data to back it up. For example, the effect of Conceal carry laws on crime rates. My professor wants the data to be on the county level with thousands of observations over years and years but that is just adding an extra layer of difficulty. Any ideas? I could use any direction for an interesting research question or useable/understandable data. I feel like this project could be easy if I have the right data and question (my prof also suggested starting with data as it could help make things easier

submitted by /u/Pleasant_Weakness_72
[link] [comments]

*In Search Of DATA* Research Project

I am in dire need of help finding a viable dataset for my research project. I am in my final semester of undergrad and have been tasked with a major research project which will soon need to be transferred into STATA but for now, I need to run basic descriptive statisitcs and come up with my hypothesis, research question, and equation. No matter what topic I bounce around I can’t seem to find data to back it up. For example, the effect of Conceal carry laws on crime rates. My professor wants the data to be on the county level with thousands of observations over years and years but that is just adding an extra layer of difficulty. Any ideas? I could use any direction for an interesting research question or useable/understandable data. I feel like this project could be easy if I have the right data and question (my prof also suggested starting with data as it could help make things easier)

submitted by /u/Pleasant_Weakness_72
[link] [comments]

Best Way To Find Resident Names From A List Of Addresses?

I have a list of addresses (including city, state, ZIP, latitude, and longitude) for a specific area, and I need to find the resident names associated with them.

I’ve already used Geocodio to get latitude and longitude, but I haven’t found a good way to pull in names. I’ve heard that services like Whitepages, Melissa Data, or Experian might work, but I’m not sure which is best or how to set it up.

Does anyone have experience with this? Ideally, I’d love a tool or API that can batch process the list. Open to paid or free solutions!

submitted by /u/Ljr1014
[link] [comments]

Movies That Were Added On Streaming Services

Hey,

I’m building my own dataset about movies that were added later on streaming services (like Netflix, Hulu, Disney+, etc). I’ve found some useful datasets in Kaggle that include the date which a specific movie was added on Netflix, for example. I need to find the dates for other movies I have in my dataset, in all other streaming services which those movies were added on. Does anyone have any idea where can I find it? When I search a specific movie in Amazon Prime, for example, I don’t find the date in which it was added on their platform.

Thanks.

submitted by /u/Porcoddio45
[link] [comments]

Paid Product Discovery Call – Dataset Procurement Protocol

Hey all!

I am building a data procurement protocol to make it easier for researchers to access proprietary datasets. We’re in the end stages of designing the UI/UX and are looking for more data points regarding pain points in the dataset procurement process. We’re offering $25 USD to anyone who can spare 15 minutes to talk about their experiences purchasing proprietary datasets. Some of the questions to expect include “Can you walk me through your typical data procurement process from identifying a need to acquiring the data?” and “Are there any specific improvements or innovations you’re hoping to implement in your data sourcing approach?”

Send me a DM if you’re interested and I’ll send you a calendly to pick a time!

submitted by /u/EmetResearch
[link] [comments]

Hello, I’m New To Datasets And Would Like To See Whether It’s Possible To Filter A Dataset From Huggingface Before Downloading It.

Hello everyone. I’m currently trying to find a more or less complete corpus of data that is completely public domain or under a free software / culture license. Something like a bundle of Wikipedia, Stack Overflow, the Gutenberg Project, and maybe some GitHub repositories for good measure. And I found RedPajama is painfully close to that, but not quite:

It includes the Common Crawl and C4 datasets, which are decidedly not completely open-source. It includes the Arxiv dataset, which might work for my purposes, but it includes both open-source and proprietary-licensed papers, so it would need filtering before I proceed. And it had to drop the Gutenberg dataset parser because of issues with it accidentally fetching copyrighted content (!!)

So, what I would like to do with RedPajama is:

Fetching Wikipedia, like usual, but also add other Wiki-projects like Wikinews and Wiktionary, and languages other than English, for completion purposes (as we’re ditching C4) Fetching more of the Stack Overflow data to compensate for the lack of C4 Fixing the Gutenberg parser so it can actually download the public-domain books from there. Alternately, download the Wikibooks dataset instead Filtering the Arxiv dataset to remove anything not under a public-domain, CC-By, or CC-By-SA license, preferably before downloading each individual paper

Is it possible to do that as a Huggingface script, or do I need to execute some manual pruning after downloading the entire RedPajama dataset instead?

submitted by /u/csolisr
[link] [comments]

Multimodal Terror Propaganda Repository Research

Looking for data on terror propaganda, Ideally multimodal e.g. social media image video audio text or others. Should be recent and ideally have a time component. The specific group behind the content is not that relevant, the more the better. There is a number of issues regarding this type of data. But as i am getting desperate i am greatful for what ever. Am looking to run some ML models for sentiment clasification tasks so need a few thousand observations. Cheers!

submitted by /u/SixMight
[link] [comments]

Looking For Options To Curate Or Download A Precurated Dataset Of Pubmed Articles On Evidence Based Drug Repositioning

To be clear, I am not looking for articles on the topic of drug repositioning, but articles that contain evidence of different drugs (for example, metformin in one case) having the potential to be repurposed for a disease other than its primary known mechanism of action or target disease (for example. metformin for Alzheimer’s). I need to be able to curate or download a dataset already curated like this. Any leads? Please help!

So far, I have found multiple ways I can curate such a database, using available API or Entrez etc. Thats good but before I put in the effort, I want to make sure there is no other way, like a dataset already curated for this purpose on kaggle or something.

For context, I am creating a RAG/LLM model that would understand connections between drugs and diseases other than the target ones.

submitted by /u/LukewarmTakesOnly
[link] [comments]

Looking For Data On Drone Delivery For Retail For A Research Project

Hey everyone,

I’m working on a research project looking into the feasibility of drones in retail delivery, and I’d really appreciate any help you could offer! My focus is mainly on a few key areas, including:

The cost-effectiveness of drone delivery How drone battery life has improved over time Changes in delivery times for drones over the past few years The number of users or corporations adopting drone delivery

That said, I’m open to any other data sets related to retail drone delivery! I’ve already looked through data sources such as AWS, Kaggle, and went through all 12 pages of Google, but I struggled to find much relevant data. The biggest challenge I’ve been facing is finding data on the costs of drone delivery and their trends, especially since many companies keep that info private.

If anyone has any data sets or knows of websites that offer this kind of data, I’d really appreciate it! Ideally, I’m looking for CSV or XLSX files, but honestly, I’m happy with any format.

Thanks so much in advance!

submitted by /u/gapple_quagsire
[link] [comments]

Just Uploaded Multiple High-Quality Datasets On Kaggle! 🚀 | IMDB, Spotify, Reddit, Air & Water Quality

Hey r/datasets

I’ve recently uploaded several diverse and high-quality datasets on Kaggle, perfect for EDA, machine learning, data visualization, and predictive modeling! If you’re looking for real-world datasets to work with, check these out:

📌 IMDB Movies Dataset 🎬

📌 Spotify Music Dataset 🎵

📌 Reddit r/todayilearned (TIL) Dataset 📜

📌 Air Quality Monitoring Dataset 🌍

📌 England Water Quality Dataset 💧

📥 Explore & Download the Datasets Here: https://www.kaggle.com/krishnanshverma/datasets

If you use any of these datasets in a project, I’d love to hear about it! Also, upvotes and feedback would be greatly appreciated to help more people discover these resources. 🚀🔥

#Kaggle #MachineLearning #DataScience #DataAnalysis #AI #BigData #OpenData

submitted by /u/krishnanshxx
[link] [comments]

Seeking Data On Children With Incarcerated Parents For A Visualization Project

Hello,

I come to you humbly! I run a small company that’s hell-bent on making a difference in the lives of children who have or had an incarcerated parent. We’re working on a project to raise awareness of the challenges these children face through data-driven storytelling and visualizations.

I’m looking for reliable datasets related to:

The number of children with incarcerated parents (preferably broken down by state or region) Demographic information (age, race, socioeconomic status) Outcomes related to education, mental health, or other relevant indicators for these children

We’ve hit multiple roadblocks in our search so far. Many schools either aren’t capturing this data because it’s not seen as a priority, or they simply don’t have the capacity to track it. If anyone knows of publicly available data sources—government reports, research studies, or anything similar—I’d be incredibly grateful for your help. This data will help inform our advocacy efforts and inspire real change.

Thanks in advance for your time and suggestions!

submitted by /u/marrthecreator
[link] [comments]

NSCH Dataset/Codebook Request 2018-2022

I’m not quite sure if this is the right place to ask for this. I’m trying to work on a project using data from the National Survey Of Children’s Health.

I was hoping someone on here would have the 2018-2022 topical data available, as well as the codebooks in SAS.

Please let me know if you’re able to share this or redirect me. They’re no longer on the website to download and I am unsure what to do.

submitted by /u/mathduckie
[link] [comments]