Category: Datatards

Here you can observe the biggest nerds in the world in their natural habitat, longing for data sets. Not that it isn’t interesting, i’m interested. Maybe they know where the chix are. But what do they need it for? World domination?

Looking For An Slop Dataset, Can Anyone Help?

Hi everyone, I am doing a personal project for a light weight way of detecting slop content (I have a super early version working in https://github.com/elalber2000/stop_slop in case you’re interested on the approach). I needed a dataset so I started searching links by hand and scrapping the content, but I would like to scale it a bit more and was wondering if maybe someone knows a dataset that could work for it. I know the term slop is not super well defined, but in this context I mean websites or text, generally AI generated (but not necessarily), that contains vague/low-effort content and is posted for seo-related objectives. I think you probably know what I mean (google is flooded with it right now), but just in case it’s not clear, this is an example of what I mean: https://visao.app/what-is-glb-file/

submitted by /u/albertus2000
[link] [comments]

Are There Any Formal References To This Dataset?

Hi all!

I’m working on a project about Multitouch Attribution Modeling using Tensor flow to predict conversion over different channels.

In the project, we are using this dataset (https://www.kaggle.com/code/hughhuyton/multitouch-attribution-modelling). However, we cannot find any formal reference (published paper or something similar) to make a proper citation. I have searched on Google a lot… really, a lot.

Does anyone know what is the origin of the data or if is it referenced somewhere?

Thanks for the help.

submitted by /u/mklsls
[link] [comments]

Looking For Comprehensive Twitter/X Posts From US Politicians

I’ve spent time searching, both online and this sub, and have found surprisingly little. I expected there to be a multiple datasets of tweets from US politicians. So far, the best I’ve found is https://www.thetrumparchive.com/ All the others are extremely limited or 5+ years old.

This seems very strange to me. This is an important record. It should exist.

I am a developer and know how to interact with APIs, but X now wants lots of money, most people don’t know how to use an API, and it’s not that helpful for going back years and years.

Am I missing something? What datasets do people use to examine the social media behavior of US politicians? Why isn’t this data readily available?

submitted by /u/pnw-steve
[link] [comments]

Looking For A Dataset To Train A Confidence Detection Model (or Advice On Building One From Scratch!)

Hey everyone! 👋 I’m working on a project to detect confidence levels in people’s speech (think job interviews, public speaking, etc.). I’m trying to rate confidence on a scale of 1-100 based on things like:

Voice characteristics (volume, pitch variation) Speaking patterns (pace, fluency, filler words) Visual cues (posture, eye contact, gestures)

I’ve been searching but haven’t found any labeled datasets specifically for confidence scoring. The closest I’ve found are emotion detection datasets, but that’s not quite what I need. Two questions:

Does anyone know of an existing dataset that scores speaker confidence? Even if it’s not public, knowing it exists would be helpful If not, what would be the best way to build this dataset?

My biggest concern is making sure the ratings are consistent and meaningful. Should I use multiple raters per video? How many samples would I need for a decent model? Really appreciate any suggestions or tips from people who’ve worked on similar problems!

Edit: This is part of a larger soft skills analysis project, so if you have experience with similar datasets (public speaking quality, interview performance etc.), I’d love to hear about those too!

submitted by /u/Fluid-Locksmith3358
[link] [comments]

Just Found This Awesome Dataset On Kaggle On Arts Auction

It’s a list of artists whose works sold for over a mil between 2018 and 2022. Proper fascinating if you’re into art, data, or both.

Why it’s cool:

Art + Data = Win: Fancy seeing which artists were raking it in? This has all the deals from Piccasso to Mark Rothko. Generate ur own arts or mix and two artistic style.

Featured Artists

Pablo Picasso (1881-1973): $2.21B total value, 245 lots sold Claude Monet (1840-1926): $1.48B total value, 89 lots sold Andy Warhol (1928-87): $1.13B total value, 136 lots sold Jean-Michel Basquiat (1960-88): $1.11B total value, 107 lots sold Gerhard Richter (b. 1932): $747.7M total value, 96 lots sold David Hockney (b. 1937): $647.2M total value, 67 lots sold Francis Bacon (1909-92): $645.5M total value, 31 lots sold Zao Wou-Ki (1920-2013): $641.3M total value, 131 lots sold Mark Rothko (1903-70): $569.6M total value, 24 lots sold

submitted by /u/Think_Huckleberry299
[link] [comments]

What Data Marketplaces Have You Used Or Know About?

Hi everyone!

I’m exploring the landscape of data marketplaces and would love to hear your experiences or recommendations.

• What data marketplaces have you used or come across?

• What stood out to you—good or bad—about their offerings or usability?

• Are there specific marketplaces you’d recommend for accessing high-quality datasets for AI, research, or business applications?

submitted by /u/Winter-Lake-589
[link] [comments]

Suggestions For Interesting Dataset For Class Project

Dear all,
I am looking for some interesting or amusing data sets that I can use for my students to do projects within a upcoming class. I have some ideas from Kaggle or the NYC open data set (the squirrel census), but I was wondering if you guys had any ideas. The audience is a semi advanced statistics class where we are going to use basic hypotheses testing up to Anova and linear regression. I just am tired of using wages and education and such.

submitted by /u/Ok_Enthusiasm428
[link] [comments]

[Dataset Request] Looking For Rural Household Economic Data For Poverty Prediction Model

I’m working on a machine learning project to predict household poverty levels in rural areas (In need the most for Cambodia dataset). I’m looking for datasets that include:

Essential features:

Household income/expenditure data Demographic information (family size, education levels, etc.) Geographic indicators (rural/urban classification) Economic indicators (employment status, assets owned) Current or historical poverty status (as target variable)

Ideal characteristics:

Recent data (preferably within the last 5-10 years) Clear documentation/data dictionary Cleaned or semi-cleaned format Country or region-level granularity Sufficient sample size for ML modeling

I’m planning to use classification techniques (Logistic Regression and XGBoost) for prediction. While I’m aware of the World Bank’s datasets, I’m interested in exploring other potential sources, especially those with more granular household-level information.

Has anyone worked with similar datasets or can point me towards reliable sources? I’m open to both public and academic databases.

Thank you in advance!

submitted by /u/Aejantou21
[link] [comments]

What Happened To / Where Is The Site That Had Huge Amounts Of Free Data For Projects?

Hi. I don’t remember the name of the site, but there was a site that had tons of tables of varying data for use in projects. I believe it was free and/or open source. If I remember correctly, it was called something like “opendata”. It’s been a few years since I’ve seen it so it might have disappeared, but I was hoping someone remembers and can point me in the right direction.

Thanks!

submitted by /u/ChargeResponsible112
[link] [comments]

I Need To Label Your Data For My Project

Hello!

I’m working on a private project involving machine learning, specifically in the area of data labeling.

Currently, my team is undergoing training in labeling and needs exposure to real datasets to understand the challenges and nuances of labeling real-world data.

We are looking for people or projects with datasets that need labeling, so we can collaborate. We’ll label your data, and the only thing we ask in return is for you to complete a simple feedback form after we finish the labeling process.

You could be part of a company, working on a personal project, or involved in any initiative—really, anything goes. All we need is data that requires labeling.

If you have a dataset (text, images, audio, video, or any other type of data) or know someone who does, please feel free to send me a DM so we can discuss the details.

submitted by /u/rafacvs
[link] [comments]

Looking For Dialect Specific Spanish Datasets

Hello everyone, I am a highschooler currently fine-tuning an LLM for translating English into accurate and specific spanish dialects, think salvadorian spanish vs cuban spanish. Its being built for warnings like hurricanes amber alerts etc… I was wondering if there were datasets that would accomplish this like conversations in salvadorian spanish?

Any help would be greatly appreciated thank you!

submitted by /u/Way2mmm
[link] [comments]

When You Guys Need To 3D Models To Use With A Game Engine For Generating Synthetic Data, Who Do You Hire And How High Do You Set Your Budgets?

I’m looking to use 3D modeled fabrications of the expected areas wherein an AR app I am developing is to be used. The app incorporates object detection, object permanence modeling, and spacial tracking. It needs to operate in a variety of conditions: clean and dirty, cluttered and no clutter, poor lighting to great lighting, and cramped to spacious. I have identified areas at my workplace that meet each of these conditions, and I want to get a rough estimate of what it would cost me to have them 3D modeled both for synthetic data generation and product testing.

submitted by /u/CurdledPotato
[link] [comments]