Category: Datatards

Here you can observe the biggest nerds in the world in their natural habitat, longing for data sets. Not that it isn’t interesting, i’m interested. Maybe they know where the chix are. But what do they need it for? World domination?

Help: Looking For Time Series Real Estate Dataset With Property Manager Info (US)

Hi everyone,

I am looking for a time series dataset of real estate properties in the United States that includes information about property managers and pricing.

Its okay if the dataset contains historical data (e.g., from 2010 to 2020) and include details such as property addresses, prices, ownership history, and the names of property managers.

If anyone knows of publicly available sources, government databases, or APIs that provide such data, I would greatly appreciate your insights. Paid sources are fine too, as long as they provide the necessary details.

Thanks in advance for your help!

submitted by /u/dank_coder
[link] [comments]

Do You Need To Prove Your Dataset Is Authentic?

I’m building a cybersecurity tool (planning to open source it soon) that helps data providers prove their data is real, accurate, and untampered.

It plugs into your database and generates cryptographic proofs and logs of all data operations, like when data was created, where it came from, and whether it was modified through life. It can also verify that queries run on the data were performed correctly.

This could be useful if you’re:

• Selling proprietary datasets and want to prove authenticity

• Sharing sensitive reports or metrics with partners

• Allowing external parties to inspect your data without exposing raw data or making copies

This isn’t about DRM or encryption. It’s about giving external users confidence in your data; without requiring them to blindly trust you.

I’d love to know:

• If you sell or license data, have buyers ever asked for proof of origin or correctness?

• If you buy datasets, do you worry about how trustworthy or fresh they are?

• Have you ever lost a deal because there was a lack of trust in the data?

Curious to know if this sounds like something useful or if it’s a solution looking for a problem.

submitted by /u/No_Telephone_9513
[link] [comments]

Any Available Datasets For Street Flood Levels?

Hi! I’m currently a 3rd year Computer Science student conducting a thesis about forecasting street floods using a machine learning model in real time. I’m currently having a hard time finding publicly available historical time-series datasets that records flood depths on urban street areas. I’ve tried Kaggle, the Google search engine for datasets, and even NASA’s Earth Data website to no avail.

I’m starting to become really worried that I might not be able to find the dataset I need to actually conduct this research. I’m planning on asking government agencies soon and other academic institutions, and see where that takes me. In the meantime, do you guys know anywhere else I could gather data for this? Do you also have any suggestions of the possible steps that I could take as a contingency plan if ever the data is actually non-existent?

Thanks!

submitted by /u/FutureFertilizer354
[link] [comments]

How To Use Multiple Languages In A Datapipeline

Was wondering if any other people here are part of teams that work with multiple different languages in a data pipeline. Eg. at my company we use some modules that are only available on R, and then run some scripts on those outputs in python. I wanted to know how teams that have this problem streamline data across multiple languages maintaining data in memory.

Are there tools that let you setup scripts in different languages to process data in a pipeline with different languages.

Mainly to be able to scale this process with tools available on the cloud.

submitted by /u/pirana04
[link] [comments]

Where Do You Source Your Data? Frustrated With Kaggle, Synthetic Data, And Costly APIs

I’m trying to build a really impressive machine learning project—something that could compete with projects from people who have actual industry experience and access to high-quality data. But I’m struggling big time with finding good data.

Most of the usual sources (Kaggle, UCI, OpenML) feel overused, and I want something unique that hasn’t already been analyzed to death. I also really dislike synthetic datasets because they don’t reflect real-world messiness—missing data, biases, or the weird patterns you only see in actual data.

The problem is, I don’t like web scraping. I know it’s technically legal in many cases, but it still feels kind of sketchy, and I’d rather not deal with potential gray areas. That leaves APIs, but it seems like every good API wants money, and I really don’t want to pay just to get access to data for a personal project.

For those of you who’ve built standout projects, where do you source your data? Are there any free APIs you’ve found useful? Any creative ways to get good datasets without scraping or paying? I’d really appreciate any advice!

submitted by /u/kobastat121987
[link] [comments]

Insights On NASA’s C-MAPSS Dataset Or ADAPT Dataset?

Hello Reddit!

In the following weeks I’ll have to start writing and conducting research for my Master’s thesis titled “Pattern recognition in industrial systems for fault detection using artificial intelligence algorithms.” My tutor has given some example datasets like Tennessee Eastman Process, CSTR, DAMADICS… But honestly I have no interest whatsoever in the field they’re in (maybe DAMADICS).

I have been searching the web for other datasets and NASA’s C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) and NASA’s ADAPT (Advanced Diagnostics and Prognostics Testbed) appear more interesting to us: windturbine lifespan, failures in spacecraft, etc.

My question is, which dataset would you recommend us focusing on? This thesis will be done in group and one of my colleagues knows a lot about machine learning since she has been working in the field quite some time, while the other colleague and I have worked with some things but not in depth. We want something that is interesting and challenging, but not excessively hard or complicated to work around.

Any insights would be appreciated! Thank you!!

submitted by /u/mustakit
[link] [comments]

What Percentage Of Bisexuals Are Trump Supporters ?

Are bisexuals statistically the least likely demographic of the LGBTQ+ to support Donald Trump and Republicans?

In my personal non statistical experience the lesbians I know all support Trump. The only lesbians I know of that don’t support Trump are Ellen Degeneres and JK Rowling. Most gay men I mean support Trump. Most trans people I meet in real life support Trump. Most gay men and trans people I see against Trump are on Reddit or content creators sites like Youtube.

But almost every bisexual I meet in real life is staunchly anti trump but there is no bisexual community like there is for gays, lesbians, trans etc so there are no bisexual content creators.

Are there any research studies done on LGBTQ trump supporting rates that can give us real numbers?

submitted by /u/lostacoshermanos
[link] [comments]

Malicious And Safe URL Dataset For ML

This dataset contains a mix of malicious and safe URLs, verified using sources like PhishTank and VirusTotal, making it ideal for training Machine Learning models. If you don’t have access to their APIs or are seeking a reliable and relevant URL dataset for ML, this is for you. This dataset will be updated daily. Cheers!

submitted by /u/TendouNoSaibA
[link] [comments]

Any Data Sets On Workers Unions Over Time?

I’m looking for data on Worker’s Unions. Number of strikes, numbers of unions, numbers of union members, numbers of contracts signed, numbers of bridge agreement/interim extension.

I’d really love to see data on union busting as well and maybe contract improvements, but I imagine those things are difficult to quantify?

I also imagine there are posts concerning this already, but I’ve already searched for ‘union’, ‘labor union’, and ‘workers union’ and haven’t come up with anything, so if there’s verbiage that I’m missing out on, feel free to chastise me for not searching so long as you tell me the terms I should have been using.

Thanks!

submitted by /u/inkblot888
[link] [comments]

What Medical Dataset Is Public For ML Research

i was trying to apply machine learning algorithm, clustering, on medical dataset to experiment if useful info comes out, but can’t find good ones.

Those in UCI repository have few rows like 300~ patient records, while many real medical papers that used ML used dataset of thousands patient records.

what medical datasets are publicly avail for ML research like this?

ps. If using dataset of 300~ patient records will be justifiable, plz also advise

submitted by /u/qmffngkdnsem
[link] [comments]

Looking For A Dataset For All London Restaurants

So I’m currently looking for a list of all restaurants in London, ideally with their M addresses.

I’ve been able to scrape a huge restaurant promotion site in the UK and pull around 7000 restaurants with this info however I’m sure I’m missing a large number of restaurants as I’m unable to find my favourite restaurants in the list.

Would anyone be able to point me in the right direction as to where I may be able to find a list like this?

submitted by /u/giveguys
[link] [comments]

Anyone Knows What Technology / Solution Was Used To Generate The Microsoft Security Incident Prediction Dataset?

So i am working on building a ML model to automate the classification of SOC environment alerts to identify the true positive ones & the false positives. The model is ready, however to be able to further test on new data, i will be needing to generate alerts similar to those that were in the training data. So if anyone has any idea what SIEM solution or EDR was used to generate these alerts, please let me know.

Microsoft Security Incident Prediction Dataset : https://www.kaggle.com/datasets/Microsoft/microsoft-security-incident-prediction?resource=download

Also are there any solutions that generate alerts with these features (OrgId, IncidentId, DetectorId, AlertId, AlertTitle, Category, Day, Id, Hour & EntityType)??

submitted by /u/Syn1ho
[link] [comments]

Any Way To Get A Set Of Seedless And Seedful Tangerine Photos?

I’m a software engineer, not super proficient in ML yet, so forgive me if my question is unrealistic.

Anyway, I want to create an app that detects whether there are seeds in a tangerine from a photo. Seedless tangerines slightly differ from seedful ones, so I believe this is somehow possible to implement. Since there is no pre-trained model for this, I’m ready to create my own, but gathering thousands of photos is an impossible mission task for me. How are tasks like this usually tackled?

submitted by /u/RoastPopatoes
[link] [comments]

Looking For A Dataset That Is Complex Enough To Do Big Data Analysis Relative To Mental Health/depression

Hello, I am in a big data class. My group is interested in doing our final project based on mental health/depression. Although, ‘big data’ will not be feasible because we are running these on our local PCs, we still need to perform big data analysis with map/reduce programs. We have been using PySpark for all of our assignments and they have been very complex assignments. Such as a friend recommendation program where you rank 10 recommendations from a very large text file that was in the format of <unique_id><list of friends>. This assignment, we had to perform multiple for loops/if statements inside of our PySpark map/reduce program which made it quite complex.

Now, we have found this dataset https://www.kaggle.com/datasets/anthonytherrien/depression-dataset that we want to use, but we don’t believe we can “wow” the professor with complex enough functions to make conclusions. Is this maybe not a good type of dataset for big data applications? We originally thought to make a depression “score” based on the given features and justify those based on how frequent/similar each unique person is.

Any ideas or datasets that you know about that would be just complex enough would be a big help. Thanks!

submitted by /u/idkwhatsgoingon4582
[link] [comments]

How To Handle Missing Values In A Dataset?

I am working on a diabetes prediction model for my project and I need help on how should I handle missing values in the smoking history column in my structured tabular dataset.

My dataset has 100,000 rows, with around 35% of rows having “No Info” for smoking history. Since smoking history has a significant impact on diabetes, this column cannot be ignored.

Other entries in this column are: “Never”, “Current”, “Not current” and “Former”

Key concerns:

Encoding: If I am encoding this column, then how should “No Info” be treated in this case? One hot encoding will lead to unneccessary high dimensionality whereas there is no clear order that I can choose between the values if I go with ordinal encoding.

Data Loss: Would dropping these rows (35%) lead to bias, or is it a valid approach?

I would appreciate your personal insights on the best approach for this since I have already searched this thing enough on the internet.

submitted by /u/shaitaanbaluck
[link] [comments]

LinkedIn Simple Dataset For Homework (how To Get?)

Hi, my teacher gave us an assignment, we need to get – how many active users by country -gender and age distributions -average users daily time on the app -percentage of the global population that uses the app. All of that in an excel or CSV. Many of my classmates had to do it with instagram, tik ton, etc. In my case it was LinkedIn, the thing is I tried to find the dataset the, only thing I could found was a statista report that I couldn’t even download. I need to put it in PowerBi so I don’t need a massive amount of data. But from what I searched in this subreddit LinkedIn API is private or I need to pay for money I don’t have.

Am not really sure on what to do, that’s why I am asking in this subreddit, where should I searched, I don’t wanna take the easy route but I spent a lot of time searching and found nothing, if there wasn’t much then u rather speak to my teacher about it. Any help would be appreciated it

submitted by /u/Jproxy122
[link] [comments]

Question For Improving Custom Floating Trash Dataset For Object Detection Model

I have a dataset of 10k images for an object detection model designed to detect and predict floating trash. This model will be deployed in marine environments, such as lakes, oceans, etc. I am trying to upgrade my dataset by gathering images from different sources and datasets. I’m wondering if adding images of trash, like plastic and glass, from non-marine environments (such as land-based or non-floating images) will affect my model’s precision. Since the model will primarily be used on a boat in water, could this introduce any potential problems? Any suggestions or tips would be greatly appreciated.

submitted by /u/Fit-Information6080
[link] [comments]

Looking For Dataset Of The Racial Wage Gap By Country

As part of a research paper, I’m currently trying to find data on the racial wage gap by country. Preferably the data will be from the at least the mid 2010’s to at least 2022, but I’d love to see anything someone can find. I’ve been looking all over the internet for it and haven’t come up with anything. Thank you!

submitted by /u/avancini12
[link] [comments]