I am appearing for a hackathon and I am looking for the datasets whcih consists of all the country’s crop and enviormental-related information. Please share necessary datasets.
submitted by /u/Mean-Pin-8271
[link] [comments]
Here you can observe the biggest nerds in the world in their natural habitat, longing for data sets. Not that it isn’t interesting, i’m interested. Maybe they know where the chix are. But what do they need it for? World domination?
I am appearing for a hackathon and I am looking for the datasets whcih consists of all the country’s crop and enviormental-related information. Please share necessary datasets.
submitted by /u/Mean-Pin-8271
[link] [comments]
it was a bunch of .txt files (containing the stories) and two xml?-files (or something) with additional metadata for the stories (title, first published, author, appeared in, rating on goodreads, rating on googlebooks etc etc) and the authors (biography, gender, name, country etc).
i remember i had to dig for it when i downloaded it like two weeks ago (just fried the laptop i saved them on, that’s why i need them again). there were some issues of the magazine Galaxy in it and a bunch of old stories: h.g. wells, asimov, de guin, and so on… i think it had a few hundred elements
if that description sounds familiar to anyone here i’d appreciate it if you could tell me where to get it again 🙂
EDIT: Christ alive, i found it: https://github.com/nschaetti/SFGram-dataset
submitted by /u/DrJotaroBigCockKujo
[link] [comments]
Hello! I’ve come across a with a lot of data form a weight loss subreddit, where people posted their starting weight, current weight, the time it took them to loseweight and their age. Now I cannot find it for the life of me. Can anyone help?
submitted by /u/ylvalloyd
[link] [comments]
Hello! Is There any kind of dataset containing recorded weapons and explosion sounds?
submitted by /u/Ablackshado
[link] [comments]
Hi folks, I’m looking for some datasets that shows comment/text contents in a multiplex social network.
Ex like: A texting family, friends, boss, B texting family, friends(maybe including A), boss…
I need the senderID, receiverID, the text content, and relation between sender and the receiver
If there’s any similar datasets plz lmk, thanks!
submitted by /u/No_Yogurt3832
[link] [comments]
I worked with someone who wanted data from one source, finished that project, enjoyed it plenty, so collected and aggregated the data from about 22 other sources. Now I have about 1M unique booze records, 430k wine records and 130k spirits record.
Wondering who i can present value to with this.
submitted by /u/makelefani
[link] [comments]
Hi all,
I am wondering if anyone knows of a data set for USA pedestrian crosswalk lights (those lights which have a red hand and counter when you should not walk and a white stick figure when you can walk). I only need USA lights however, all I can find are datasets for China or UK. Any help appreciated.
submitted by /u/aadiman23
[link] [comments]
Hi everyone,
not sure if this is the correct thread, but hoping it is. So long story short, I am trying to compile a database of every indian politician (I have a list of them by name/party which ive imported into excel). I need to include their date of birth, date of death. Many politicians have wikipedia pages so currently I am manually going through each politician, searching them up and then entering their details into the database.
Would there be any faster way to do this? I am doing this for a scientific paper so i need it done asap but this method seems like it would take forever
submitted by /u/Aggravating_Hope2390
[link] [comments]
Hello!
I’m looking for a dataset about occupations by race and ethnicity, gender and age (or any related demographic information) in the state of California. If there is a national dataset that would work too.
Thank you!
submitted by /u/htxastrowrld
[link] [comments]
Hello,
I’m looking for air traffic data for a personal project. I’d like to be able find the number of direct flights (ideally passengers too, but this is probably too granular) between two US airports. The way I’m envisioning it, the data for a given period of time would look something like the matrix below, where departure airports are column headers, destination airports are row headers, and the values are the number of flights (or passengers):
Airport ATL DFW DEN … ATL XXX 12 8 … DFW 11 XXX 10 … DEN 9 10 XXX … … … … … XXX
I know the data needed to create something like this must exist somewhere, but the closer to the end product displayed above the better. Thanks!
submitted by /u/dobby_bodd
[link] [comments]
Hi,
I am a seasoned market researcher and got intrigued with data science and machine learning since most of my job is about dealing with data. I am currently pursuing my MSc in Data Science. Before this we were provided with datasets to work on. It was initially a struggle to define my own use case based on the data that was shared, however, I was able to deliver with average results.
However, for my next coursework we should be using our own datasets which should be supervised learning in nature and they cannot be from Kaggle or UCI (we lose 30 points if we use any of these sources for our datasets. I have spent about a week to look for datasets and I am a bit confused and also unable to understand which dataset to use or what kind of use cases should I look at. I did explore data.gov but I kind of just freeze because I am unable to understand what use case I can create of the database. I can’t use clustering problem because that would be unsupervised in nature.
I tried a couple of regional sites for web scraping – use cases tried were second hand car and predicting their price and rental price prediction based on area selected. However the websites did not allow web scraping and I would like to respect that.
Would you know any publicly available datasets that I can potentially explore for my supervised machine learning coursework?
Do let me know if you have any idead that I can explore and thanks in advance.
submitted by /u/jknotra
[link] [comments]
Hello everyone,
I am starting a new job for a company, my role is Data Specialist and I will be responsible for working on helping the team with Data migration.
The task is to do data migration from legacy system to a modern data structure on SharePoint. I am very much aware of the steps involved in this process, however, I have been out of touch with the tools and techniques as last time I studied Python, SSIS, visualization and Excel tools was a few years back and I think it will be difficult for me to contribute immediately as I join them. I am starting my work next week. I wanted to ask you professionals if I would have a hard time at my work with the skills I don’t possess right now and what are the steps I need to take to make sure my employer can count on me going forward.
PS: This is my first day working for an IT company and I have no idea how an IT project works.
Thanks!
submitted by /u/shanke_y8
[link] [comments]
Hello everyone!
So I wanted to do an EDA project on SQL but I don’t know where to start. I would prefer to work with an HR related dataset as I’m more familiar with HR and find the data already interesting.
Can someone point to a good dataset for this? Thanks!
submitted by /u/cam171811
[link] [comments]
Hi,
I’m working on my thesis and I wonder if anyone has this data or can lead me to relevant governmental web page that contains this?
Cheers!
submitted by /u/rolanddes1
[link] [comments]
I’m looking for a dataset that contains a large amount of features (100+) for clustering purposes. I’m applying sparse clustering, which means that my method removes features that are not important for the clusters. Either a data set that only contains categorical variables, or a mix of numerical and categorical variables is fine.
Does anyone have any ideas?
submitted by /u/y_zh
[link] [comments]
Where can I find labelled spam tweets dataset that has at least 50000 rows. I’ve searched everywhere, but there are very few datasets available, and those datasets are way too small. It is impossible to use the API to get the tweets without paying 🙁
I am in urgent need of one, as I will need it for my MSc dissertation, any leads are greatly appreciated.
Thanks
submitted by /u/kingsterkrish
[link] [comments]
I need a dataset(s) that contains the recordings of one person: -temperature -heart rate -oxygen saturation.
For 24 hours with a timestamp of 1 min
submitted by /u/Lili23data
[link] [comments]
I think this is an interesting dataset that I generated from ChatGPT, but I am not sure how to generate visuals for it. Does anyone have any suggestions?
Height (Wealth Level) Percentage of Population Rough Number of People Rough Wealth Range 1 inch (Poverty Level) 10.5% ~34.8 million $0 – $10,000 5 feet (Median Wealth) 50% ~165.5 million $10,000 – $100,000 6 feet (Affluent) 25% ~82.75 million $100,000 – $1,000,000 10 feet (Wealthy) 10% ~33.1 million $1,000,000 – $10,000,000 100 feet (Ultra-Wealthy) 1% ~3.31 million $10,000,000 – $1,000,000,000 1000 feet (Billionaire Class) 0.0002% ~660 individuals $1,000,000,000 and above
submitted by /u/eagle_eye_johnson
[link] [comments]
Looking to put together a searchable website where users can search a product and see the price / quantity deferential over the past 2-3 years. Any help pointing me in the right direction would be greatly appreciated.
submitted by /u/mortalhal
[link] [comments]
i need a small dataset of vehicles detection have at least 6 classes contain ambulance class for my project “smart traffic control” , I am use YOLO v8
submitted by /u/ibrahim-elsadat
[link] [comments]
Anyone have another source to download the REDD dataset? The original source http://redd.csail.mit.edu/, is not working and cant access. We badly need the full dataset. Thank you!
submitted by /u/luna_anabanana
[link] [comments]
Hello, everyone! Does anyone here has data set for mushroom yield production that includes temperature and humidity data? We need at least 1,500 data for our simulation as part of our capstone project. Thank you.
submitted by /u/Ill-Moose4794
[link] [comments]
I would really appreciate feedback on a version control for tabular datasets I am building, the Data Manager.
Main characteristics:
Like DVC and Git LFS, integrates with Git itself. Like DVC and Git LFS, can store large files on AWS S3 and link them in Git via an identifier. Unlike DVC and Git LFS, calculates and commits diffs only, at row, column, and cell level. For append scenarios, the commit will include new data only; for edits and deletes, a small diff is committed accordingly. With DVC and Git LFS, the entire dataset is committed again, instead: committing 1 MB of new data 1000 times to a 1 GB dataset yields more than 1 TB in DVC (a dataset that increases linearly in size between 1 GB and 2 GB, committed 1000 times, results in a repository of ~1.5 TB), whereas it sums to 2 GB (1 GB original dataset, plus 1000 times 1 MB changes) with the Data Manager. Unlike DVC and Git LFS, the diffs for each commit remain visible directly in Git. Unlike DVC and Git LFS, the Data Manager allows committing changes to datasets without full checkouts on localhost. You check out kilobytes and can append data to a dataset in a repository of hundreds of gigabytes. The changes on a no-full-checkout branch will need to be merged into another branch (on a machine that does operate with full checkouts, instead) to be validated, e.g., against adding a primary key that already exists. Since the repositories will contain diff histories, snapshots of the datasets at a certain commit have to be recreated to be deployable. These can be automatically uploaded to S3 and labeled after the commit hash, via the Data Manager.
Links:
https://news.ycombinator.com/item?id=35930895 [no full checkout] https://youtu.be/BxvVdB4-Aqc https://news.ycombinator.com/item?id=35806843 [general intro] https://youtu.be/J0L8-uUVayM
This paradigm enables hibernating or cleaning up history on S3 for old datasets, if these are deleted in Git and snapshots of earlier commits are no longer needed. Individual data entries can also be removed for GDPR compliance using versioning on S3 objects, orthogonal to git.
I built the Data Manager for a pain point I was experiencing: it was impossible to (1) uniquely identify and (2) make available behind an API multiple versions of a collection of datasets and config parameters, (3) without overburdening HDDs due to small, but frequent changes to any of the datasets in the repo and (4) while being able to see the diffs in git for each commit in order to enable collaborative discussions and reverting or further editing if necessary.
Some background: I am building natural language AI algorithms (a) easily retrainable on editable training datasets, meaning changes or deletions in the training data are reflected fast, without traces of past training and without retraining the entire language model (sounds impossible), and (b) that explain decisions back to individual training data.
I look forward to constructive feedback and suggestions!
submitted by /u/Usual-Maize1175
[link] [comments]
I am creating Career Recommender for my Computer Science Project and I am looking for a dataset that contains information about all the different Careers/Jobs and stats about them like salary, field, and Education needed for it. Where can I find a dataset which would include these are at least some?
submitted by /u/parrotpo
[link] [comments]
I am working on a project to collect data from Different sources (distributors, retail stores, etc.) thru different approaches (ftp, api, scrapping, excel, etc.). I would like to consolidate all the information and create dynamic reports. I would like to add all the offers and discounts suggested by these various vendors.
How do I get all this data? Is there a data provider who can provide the data? I would like to start with IT hardware and IT Electronic Consumers goods.
Any help is highly appreciated. TIA
submitted by /u/BeGood9170
[link] [comments]