Looking For Elementary Or Secondary School Data In China.

I’m looking for school data for any province or municipality in China. Ideally, school-level variables including achievement, enrolment, or SES.

submitted by /u/Dr_Mokiki
[link] [comments]

0

Looking For Haitian Creole Voice Dataset

I’m looking a Haitian Creole audio dataset to develop a translation tool to serve the Haitian migrants worldwide communities. I found some but they’re not enough to create something robust for accuracy and good pronunciation.

Please help!

submitted by /u/Objective-Mood-6467
[link] [comments]

0

Looking For Ad Detection In Text Datasets

I have a bunch of audio and video files which have ads in them. My plan was to get transcripts of these files (maybe using whisper but not confirmed yet) and then detect which timestamps have ads on them. Anyone know any datasets that could help with this?

submitted by /u/Captainphilipp21
[link] [comments]

0

[Dataset Request] Looking For Rural Household Economic Data For Poverty Prediction Model

I’m working on a machine learning project to predict household poverty levels in rural areas (In need the most for Cambodia dataset). I’m looking for datasets that include:

Essential features:

Household income/expenditure data Demographic information (family size, education levels, etc.) Geographic indicators (rural/urban classification) Economic indicators (employment status, assets owned) Current or historical poverty status (as target variable)

Ideal characteristics:

Recent data (preferably within the last 5-10 years) Clear documentation/data dictionary Cleaned or semi-cleaned format Country or region-level granularity Sufficient sample size for ML modeling

I’m planning to use classification techniques (Logistic Regression and XGBoost) for prediction. While I’m aware of the World Bank’s datasets, I’m interested in exploring other potential sources, especially those with more granular household-level information.

Has anyone worked with similar datasets or can point me towards reliable sources? I’m open to both public and academic databases.

Thank you in advance!

submitted by /u/Aejantou21
[link] [comments]

0

Medical Dataset Sources Required …

I wanted to train some models and wanted to try maybe retina scans or x-rays or anything but couldn’t find any good sources for it besides kaggle. Does anyone have any other good sources I can use

submitted by /u/Boboflip27
[link] [comments]

0

Seeking Data On Areas Destroyed Or Number Of Martyrs From The Last War On Gaza Strip

I have portfolio project at correlation one “data analysis” program , and i decided it to be related to the last war of Gaza, I need resources if any could provide to me.

submitted by /u/Forsaken-Brilliant81
[link] [comments]

0

What Happened To / Where Is The Site That Had Huge Amounts Of Free Data For Projects?

Hi. I don’t remember the name of the site, but there was a site that had tons of tables of varying data for use in projects. I believe it was free and/or open source. If I remember correctly, it was called something like “opendata”. It’s been a few years since I’ve seen it so it might have disappeared, but I was hoping someone remembers and can point me in the right direction.

Thanks!

submitted by /u/ChargeResponsible112
[link] [comments]

0

Looking For A PC Game System Requirements Dataset

Hey everyone, I’m searching for a dataset that contains system requirements for PC games. If anyone knows where I can find such a dataset or has a link to one, I’d greatly appreciate it! 🙂

submitted by /u/NarrowGiraffe6444
[link] [comments]

0

Public Domain Image Archive. Find Images You Can Use

submitted by /u/cavedave
[link] [comments]

0

I Need To Label Your Data For My Project

Hello!

I’m working on a private project involving machine learning, specifically in the area of data labeling.

Currently, my team is undergoing training in labeling and needs exposure to real datasets to understand the challenges and nuances of labeling real-world data.

We are looking for people or projects with datasets that need labeling, so we can collaborate. We’ll label your data, and the only thing we ask in return is for you to complete a simple feedback form after we finish the labeling process.

You could be part of a company, working on a personal project, or involved in any initiative—really, anything goes. All we need is data that requires labeling.

If you have a dataset (text, images, audio, video, or any other type of data) or know someone who does, please feel free to send me a DM so we can discuss the details.

submitted by /u/rafacvs
[link] [comments]

0

The Best Tacit Knowledge Videos On Every Subject

submitted by /u/cavedave
[link] [comments]

0

Open Source Credit Risk With Telco Dataset

I am looking to develop a loan approval model solely based on applicant mobile data (make, model, specs etc.). Can anyone suggest an online data source that contains device info in addition to credit bureau and finance data? (have looked into openML, UCI and Kaggle with no luck). Thanks!

submitted by /u/no_bullshit_sherlock
[link] [comments]

0

Looking For Dialect Specific Spanish Datasets

Hello everyone, I am a highschooler currently fine-tuning an LLM for translating English into accurate and specific spanish dialects, think salvadorian spanish vs cuban spanish. Its being built for warnings like hurricanes amber alerts etc… I was wondering if there were datasets that would accomplish this like conversations in salvadorian spanish?

Any help would be greatly appreciated thank you!

submitted by /u/Way2mmm
[link] [comments]

0

Sell Data To China – Individual Level

how does one go about selling there personal data to china?? Like on an individual level?? If congress is gonna ban tik tok I want to personally send that shit to the chinese communist party out of spite…. #freegaza

submitted by /u/basedghoul
[link] [comments]

0

How Do Sites Like Character.AI, Replika And Candy.ai Get Datasets For Their Thousands Of Characters???

I am building something similar as a project and I don’t understand how to power the characters with different personalities. chatGPT suggested that fine tuning models are each character would be the way but how should i do that if I have no datasets or anything to do that, guide me to the right direction, thanks

submitted by /u/smallchindude
[link] [comments]

0

When You Guys Need To 3D Models To Use With A Game Engine For Generating Synthetic Data, Who Do You Hire And How High Do You Set Your Budgets?

I’m looking to use 3D modeled fabrications of the expected areas wherein an AR app I am developing is to be used. The app incorporates object detection, object permanence modeling, and spacial tracking. It needs to operate in a variety of conditions: clean and dirty, cluttered and no clutter, poor lighting to great lighting, and cramped to spacious. I have identified areas at my workplace that meet each of these conditions, and I want to get a rough estimate of what it would cost me to have them 3D modeled both for synthetic data generation and product testing.

submitted by /u/CurdledPotato
[link] [comments]

0

[Dataset] 19,762 Garbage Images In 10 Classes For AI And Sustainability

Hi everyone,

I’ve just released a new version of the Garbage Classification V2 Dataset on Kaggle. This dataset contains 19,762 high-quality images categorized into 10 classes of common waste items:

Metal: 1020 Glass: 3061 Biological: 997 Paper: 1680 Battery: 944 Trash: 947 Cardboard: 1825 Shoes: 1977 Clothes: 5327 Plastic: 1984

Key Features:

Diverse Categories: Covers common household waste items. Balanced Distribution: Suitable for robust ML model training. Real-World Applications: Ideal for AI-based waste management, recycling programs, and educational tools.

🔗 Dataset Link: Garbage Classification V2

This dataset has already been featured in the research paper, “Managing Household Waste Through Transfer Learning.” Let me know how you’d use this in your projects or research. Your feedback is always welcome!

submitted by /u/Downtown_Bag8166
[link] [comments]

0

GitHub – Adverse-media-dataset: Weekly Free Adverse Media News Datasets From Global News Sites

submitted by /u/rangeva
[link] [comments]

0

Need Images Of Human Arms For Dataset

Hey! I am in the process of creating a dataset for detecting human skin/arms from a close range.

I have gathered about 500 images and drawn polygons around the arms from a close range, I did this by taking photos of my own arms and asking my friends to take similar pictures but I think I still need about 500 more images. Is there anyway I could get more similar images quickly?

Open to posting job ads, is there a place to ask for images of this sort?

I have attached an imgur of images im looking for. thanks for reading!

Notes: I have already scowered all the stock images on google, as well as gone through every “arm” related dataset on roboflow

https://imgur.com/a/arm-XZGHgTP – Here are reference image

submitted by /u/blur69xd
[link] [comments]

0

[Dataset] Testing The “Pinnacle EV Betting” Theory: FanDuel Vs Pinnacle NFL Line Accuracy (2020-2023)

Dataset Referenced: https://github.com/bentodd1/FanDuelVsPinnacle/blob/master/line_comparison.csv

Background: While building smartbet.name, I noticed many betting sites claim you can do EV betting by following Pinnacle’s lines. I decided to test this by comparing Pinnacle and FanDuel NFL lines, with surprising results.

Key Findings:

Dataset: 1,039 NFL games (2020-2023) Lines from both books captured week before games FanDuel showed better predictive accuracy

Results Breakdown:

Line Accuracy: Identical predictions: 457 games (43.98%) FanDuel more accurate: 302 games (29.07%) Pinnacle more accurate: 280 games (26.95%) Average Absolute Error: Pinnacle: 9.51 points FanDuel: 9.05 points Average Hours Before Game: Pinnacle: 88.1 hours FanDuel: 58.0 hours

Dataset Access:

Full Dataset: line_comparison.csv Analysis Code: Jupyter Notebook

Methodology: The exact analysis can be seen in the Jupyter notebook. I created the database while using smartbet.name .

These findings challenge conventional wisdom about Pinnacle’s supposed edge in market efficiency.

submitted by /u/bentodd1
[link] [comments]

0

Help Finding Data: Measure Of Tourism

Hi guys, I’m doing my dissertation on the effect of precipitation on different factors of tourism within Ireland. I’m really struggling to find the dataset I need. I’m looking for any sort of measure of tourism eg. Visitor numbers, hotel occupancy, estimated tourist expenditure (anything at this point) that spans about 10 years, is monthly data, and also a regional scope of Ireland (Dublin, west coast, east coast ect.) I’ve been searching for a while now and have a few datasets but nothing perfect. Please let me know if you have any tips or even know of a dataset which may help. Thanks!

submitted by /u/MessBig6240
[link] [comments]

0

Looking For Prescription Data Of Medicine In Different Countries

The Netherlands publishes the amount of each drug prescribed and dispensed in a certain time periode (https://www.gipdatabank.nl/). For a small comparison in which drugs are used in which country I need the same data from other countries (at least the G20 countries).

Had some rough battles with the NHS site for example, but can’t really find the data in the same way, organized by ATC. Any pointers on where to look?

submitted by /u/Koopabro
[link] [comments]

0

Finding Datasets Of Images Paired With Air Quality

I’m trying to train a vision classifier to estimate air quality just from images.

Currently I’m scraping public webcams and using nearby air quality. But it’s not diverse enough. I only got two webcams with bad air quality and they’re all in China.

Are there any other good ways to find this?

submitted by /u/PathonScript
[link] [comments]

0

How Is The Research Community Dealing With Twitter Banning Scapping?

I am fairly new to the NLP field. Most of the papers in the literature perform text analysis on twitter data. Now that twitter has clamped down on scraping, how can one get the twitter post data? How is the research community dealing with it?

submitted by /u/Comprehensive-Ad1072
[link] [comments]

0

Spotify Dataset Cleaning JSON Files Help

Hello all so as the title says I downloaded a dataset from my spotify thats my listening history. I want to clean the data to eventually bring them into sql so i can create tables etc any tips first time working with JSON and its been a couple months off of coding.

submitted by /u/Far-Upstairs8318
[link] [comments]

0

Just Find A Open Source Fitness Dataset

submitted by /u/new7dev
[link] [comments]

0

Spotify Data On Amount Of Times A Link To A Song Has Been Copied And Or Shared?

I’m currently working on a project exploring social herding in music consumption and was wondering whether there is any data on this. Any data on anything like “referral links” would make this project much easier. Very grateful for any and all input / help, thanks in advance!

submitted by /u/carlmakesmusic
[link] [comments]

0

High Resolution Heat Pump Harmonics Data

submitted by /u/Sudden-Host-642
[link] [comments]

0

Biomedical Reasoning 10k Synthetic Dataset – Experimented With Data Mixes Until This One. 1.1B TinyLlama Beats GPT 4o Mini On PubMedQA With This

submitted by /u/Classic_Eggplant8827
[link] [comments]

0

Choosing One Financial Institution Over Other Ones

Hi! I would appreciate any help in advance! The question we like to answer is:

why consumers choose one financial institution over another for mortgage loans. Factors to consider include interest rates, fees, reputation, trust, loan terms, customer service, approval speed, product offerings, convenience, recommendations, financial stability, and special offers.

Therefore I need datasets that explicitly have consumers side, whether or not choosing one institution. One I found interesting is HDMA datasets that has one class of applicants who are approved for a loan but did not accepted the loan. It’s interesting, but has not much new to say or significantly different factors than other ones like those who accepted the loan or got denied. I was wondering if there are other datasets that might have consumers side of view showing factors that impact consumers decisions? Anything that might expand my perspective, basically. Thanks!

submitted by /u/Responsible-Ice-874
[link] [comments]

0

Category: Datatards

Looking For Elementary Or Secondary School Data In China.

Looking For Haitian Creole Voice Dataset

Looking For Ad Detection In Text Datasets

[Dataset Request] Looking For Rural Household Economic Data For Poverty Prediction Model

Medical Dataset Sources Required …

Seeking Data On Areas Destroyed Or Number Of Martyrs From The Last War On Gaza Strip

What Happened To / Where Is The Site That Had Huge Amounts Of Free Data For Projects?

Looking For A PC Game System Requirements Dataset

Public Domain Image Archive. Find Images You Can Use

I Need To Label Your Data For My Project

The Best Tacit Knowledge Videos On Every Subject

Open Source Credit Risk With Telco Dataset

Looking For Dialect Specific Spanish Datasets

Sell Data To China – Individual Level

How Do Sites Like Character.AI, Replika And Candy.ai Get Datasets For Their Thousands Of Characters???

When You Guys Need To 3D Models To Use With A Game Engine For Generating Synthetic Data, Who Do You Hire And How High Do You Set Your Budgets?

[Dataset] 19,762 Garbage Images In 10 Classes For AI And Sustainability

Key Features:

GitHub – Adverse-media-dataset: Weekly Free Adverse Media News Datasets From Global News Sites

Need Images Of Human Arms For Dataset

[Dataset] Testing The “Pinnacle EV Betting” Theory: FanDuel Vs Pinnacle NFL Line Accuracy (2020-2023)

Help Finding Data: Measure Of Tourism

Looking For Prescription Data Of Medicine In Different Countries

Finding Datasets Of Images Paired With Air Quality

How Is The Research Community Dealing With Twitter Banning Scapping?

Spotify Dataset Cleaning JSON Files Help

Just Find A Open Source Fitness Dataset

Spotify Data On Amount Of Times A Link To A Song Has Been Copied And Or Shared?

High Resolution Heat Pump Harmonics Data

Biomedical Reasoning 10k Synthetic Dataset – Experimented With Data Mixes Until This One. 1.1B TinyLlama Beats GPT 4o Mini On PubMedQA With This

Choosing One Financial Institution Over Other Ones

Recent Posts

Recent Comments

18+ Content

Key Features:

Recent Posts

Recent Comments