I would require the full video ideally to download not the features
Ideally internet shared, compressed etc.
already trying out webvid so suggest others
thank you
submitted by /u/ChestFree776
[link] [comments]
Here you can observe the biggest nerds in the world in their natural habitat, longing for data sets. Not that it isn’t interesting, i’m interested. Maybe they know where the chix are. But what do they need it for? World domination?
I would require the full video ideally to download not the features
Ideally internet shared, compressed etc.
already trying out webvid so suggest others
thank you
submitted by /u/ChestFree776
[link] [comments]
Need to add these to a project for my masters but it seems impossible to find – would anyone have any idea where?
submitted by /u/Puzzled_Potato_931
[link] [comments]
Hi everyone,
I’m looking for advice on where to find or purchase a comprehensive, up-to-date global dataset of tech companies and startups for 2026.
I need a global reach (US, EU, Asia) and specifically require datasets that include:
• Company Name
• Official Website URL
• Verified Business Email
I want to avoid outdated lists and “dead” websites from previous years. Does anyone know of reliable providers, directories, or platforms that offer high-quality global exports for this year?
Any recommendations for tools or marketplaces that specialize in recently updated business data would be greatly appreciated.
Thanks!
submitted by /u/Embarrassed_Fig_566
[link] [comments]
Hi guys, I’m conducting research to support awareness for men’s mental health. If any guy who is 18 or older can fill out my 2 minute survey, I would greatly appreciate it. Thank you!!
submitted by /u/bkyle27
[link] [comments]
I’ve been looking for large-scale RGB-D datasets that actually keep the naturally occurring depth holes from consumer sensors instead of filtering them out or only providing clean rendered ground truth. Most public RGB-D datasets (ScanNet++, Hypersim, etc.) either avoid challenging materials or give you near-perfect depth, which is great for some tasks but useless if you’re trying to train models that handle real sensor failures on glass, mirrors, metallic surfaces, etc.
Recently came across the data released alongside the LingBot-Depth paper (“Masked Depth Modeling for Spatial Perception”, arXiv:2601.17895). They open-sourced 3M RGB-D pairs (2M real + 1M synthetic) specifically curated to preserve the missing depth patterns you get from actual hardware.
What’s in the dataset:
| Split | Samples | Source | Notes |
|---|---|---|---|
| LingBot-Depth-R | 2M | Real captures (Orbbec Gemini, Intel RealSense, ZED) | Homes, offices, gyms, lobbies, outdoor scenes. Pseudo GT from stereo IR matching with left-right consistency check |
| LingBot-Depth-S | 1M | Blender renders + SGM stereo | 442 indoor scenes, includes speckle-pattern stereo pairs processed through semi-global matching to simulate real sensor artifacts |
| Combined training set | ~10M | Above + 7 open-source datasets (ClearGrasp, Hypersim, ARKitScenes, TartanAir, ScanNet++, Taskonomy, ADT) | Open-source splits use artificial corruption + random masking |
Each real sample includes synchronized RGB, raw sensor depth (with natural holes), and stereo IR pairs. The synthetic samples include RGB, perfect rendered depth, stereo pairs with speckle patterns, GT disparity, and simulated sensor depth via SGM. Resolution is 960×1280 for the synthetic branch.
The part I found most interesting from a data perspective is the mask ratio distribution. Their synthetic data (processed through open-source SGM) actually has more missing measurements than the real captures, which makes sense since real cameras use proprietary post-processing to fill some holes. They provide the raw mask ratios so you can filter by corruption severity.
The scene diversity table in the paper covers 20+ environment categories: residential spaces of various sizes, offices, classrooms, labs, retail stores, restaurants, gyms, hospitals, museums, parking garages, elevator interiors, and outdoor environments. Each category is roughly 1.7% to 10.2% of the real data.
Links:
HuggingFace: https://huggingface.co/robbyant/lingbot-depth
GitHub: https://github.com/robbyant/lingbot-depth
Paper: https://arxiv.org/abs/2601.17895
The capture rig is a 3D-printed modular mount that holds different consumer RGB-D cameras on one side and a portable PC on the other. They mention deploying multiple rigs simultaneously to scale collection, which is a neat approach for anyone trying to build similar pipelines.
I’m curious about a few things from anyone who’s worked with similar data:
submitted by /u/Electrical-Shape-266
[link] [comments]
I have compiled a structured dataset covering every league match in the Premier League, La Liga, Bundesliga, Serie A, and Ligue 1 from the 2015/16 season to the present.
• Format: Weekly JSON/XML files (one file per league per game-week)
• Player-level detail per appearance: minutes played (start/end), goals, assists, shots, shots on target, saves, fouls committed/drawn, yellow/red cards, penalties (scored/missed/saved/conceded), own goals
• Approximate volume: 1,860 week-files (~18,000 matches, ~550,000 player records)
The dataset was originally created for internal analysis. I am now considering offering the complete archive as a one-time ZIP download.
I am assessing whether there is genuine interest from researchers, analysts, modelers, or others working with football data.
If this type of dataset would be useful for your work (academic, modeling, fantasy, analytics, etc.), please reply with any thoughts on format preferences, coverage priorities, or price expectations.
I can share a small sample week file via DM or comment if helpful to evaluate the structure.
submitted by /u/Specialist-Hand6171
[link] [comments]
Most ESG datasets rely on corporate self-disclosures — companies grading their own homework. This dataset takes a fundamentally different approach. Every score is derived from adversarial sources that companies cannot control: court filings, regulatory fines, investigative journalism, and NGO reports.
The dataset contains integrity scores for all S&P 500 companies, scored across 11 ethical dimensions on a -100 to +100 scale, where -100 represents the worst possible conduct and +100 represents industry-leading ethical performance.
Each row represents one S&P 500 company. The key fields include:
Company information: ticker symbol, company name, stock exchange, industry sector (ISIC classification)
Overall rating: Categorical assessment (Excellent, Good, Mixed, Bad, Very Bad)
11 dimension scores (-100 to +100):
planet_friendly_business — emissions, pollution, environmental stewardship
honest_fair_business — transparency, anti-corruption, fair practices
no_war_no_weapons — arms industry involvement, conflict zone exposure
fair_pay_worker_respect — labour rights, wages, working conditions
better_health_for_all — public health impact, product safety
safe_smart_tech — data privacy, AI ethics, technology safety
kind_to_animals — animal welfare, testing practices
respect_cultures_communities — indigenous rights, community impact
fair_money_economic_opportunity — financial inclusion, economic equity
fair_trade_ethical_sourcing — supply chain ethics, sourcing practices
zero_waste_sustainable_products — circular economy, waste reduction
Traditional ESG providers (MSCI, Sustainalytics, Morningstar) rely heavily on corporate sustainability reports — documents written by the companies themselves. This creates an inherent conflict of interest where companies with better PR departments score higher, regardless of actual conduct.
This dataset is built using NLP analysis of 50,000+ source documents including:
Court records and legal proceedings
Regulatory enforcement actions and fines
Investigative journalism from local and international outlets
Reports from NGOs, watchdogs, and advocacy organisations
The result is 11 independent scores that reflect what external evidence says about a company, not what the company says about itself.
Alternative ESG analysis — compare these scores against traditional ESG ratings to find discrepancies
Ethical portfolio screening — identify S&P 500 holdings with poor conduct in specific dimensions
Factor research — explore correlations between ethical conduct and financial performance
Sector analysis — compare industries across all 11 dimensions
ML/NLP research — use as labelled data for corporate ethics classification tasks
ESG score comparison — benchmark against MSCI, Sustainalytics, or Refinitiv scores
Scores are generated by Mashini Investments using AI-driven analysis of adversarial source documents.
Each company is evaluated against detailed KPIs within each of the 11 dimensions.
– 500 companies — S&P 500 constituents
– 11 dimensions — 5,533 individual scores
– Score range — -100 (worst) to +100 (best)
CC BY-NC-SA 4.0 licence.
submitted by /u/RevolutionaryGate742
[link] [comments]
I’d love to use the polymarket data of the US 2024 presidential election. Is there any way I can get an export of timestamped data (hourly level, last 3-4 days before event)? I’m in Europe where Polymarket can only be accessed through VPN, maybe I can even find a dataset somewhere. Thanks a lot!
submitted by /u/selammeister
[link] [comments]
I’ve recently started collecting an early-stage, fully anonymous dataset
showing aggregated stress scores by country and state.
The data is derived from on-device computations and shared only as a single
daily score per region (no raw signals, no personal data).
Coverage is still limited, but the dataset is growing gradually.
Sharing here mainly to document the dataset and gather early feedback.
Public overview and weekly summaries are available here:
submitted by /u/maxstrok
[link] [comments]
Hi everyone,
I’m offering a large-scale EU Amazon product intelligence dataset with 4 million+ entries, continuously updated.
The dataset is primarily focused on high resale-value products (electronics, lighting, branded goods, durable products, etc.), making it especially useful for arbitrage, pricing analysis, and market research. US Amazon data will be added shortly.
What’s included:
Dataset characteristics:
Delivery & Format:
Who this is for:
Sample & Demo:
A small sample (10–50 records) is available on request so you can review structure and data quality before purchasing.
Pricing & Payment:
If this sounds useful, feel free to DM me — happy to share a sample or discuss a custom extract.
submitted by /u/Fun_Internal1460
[link] [comments]
Hi everyone,
I’m looking for the Yahoo S5 KPI Anomaly Detection dataset for research purposes.
If anyone has a link or can share it, I’d really appreciate it!
Thanks in advance.
submitted by /u/PrestigiousHeight76
[link] [comments]
Hello everyone, I’d like to share a high-fidelity synthetic dataset I developed for research and testing purposes.
Please note that the link is to my personal store on Gumroad, where the dataset is available for sale.
Technical Details:
I generated 1,000,000 records based on diabetes health indicators (original source BRFSS 2015) using Gaussian Copula models (SDV library).
• Privacy: The data is 100% synthetic. No risk of re-identification, ideal for development environments requiring GDPR or HIPAA compliance.
• Quality: The statistical correlations between risk factors (BMI, hypertension, smoking) and diabetes diagnosis were accurately preserved.
• Uses: Perfect for training machine learning models, benchmarking databases, or stress-testing healthcare applications.
Link to the dataset: https://borghimuse.gumroad.com/l/xmxal
Feedback and questions about the methodology are welcome!
submitted by /u/Same_Asparagus_1979
[link] [comments]
I need a data set around the immigrant health paradox. Specifically one that analyzes the shifts in immigrants health the longer they stay in US by age group. #dataset#data analysis
submitted by /u/Individual_Type4123
[link] [comments]
Hi everyone,
I’m working on building a retail data analysis portfolio project and wanted to ask if anyone here has worked on or built a good retail analysis project that they’d be willing to share or let me refer to.
I’m mainly looking for project ideas, problem statements, datasets, or dashboards that reflect real-world retail use cases (sales analysis, customer behavior, inventory, forecasting, etc.).
Any links, GitHub repos, or brief descriptions would be really helpful.
Thank you in advance.. I would really appreciate your time and help! 😊
submitted by /u/msfarahs
[link] [comments]
LLM agent benchmarks like τ-bench ask what agents can do. Real deployment asks something harder: do they know when they shouldn’t act?
CAR-bench (https://arxiv.org/abs/2601.22027), a benchmark for automotive voice assistants with domain-specific policies, evaluates three critical LLM Agent capabilities:
1️⃣ Can they complete multi-step requests?
2️⃣ Do they admit limits—or fabricate capabilities?
3️⃣ Do they clarify ambiguity—or just guess?
Three targeted task types:
→ Base (100 tasks): Multi-step task completion
→ Hallucination (90 tasks): Admit limits vs. fabricate
→ Disambiguation (50 tasks): Clarify vs. guess
tested in a realistic evaluation sandbox:
58 tools · 19 domain policies · 48 cities · 130K POIs · 1.7M routes · multi-turn interactions.
What was found: Completion over compliance.
SOTA model (Claude-Opus-4.5): only 52% consistent success.
Hallucination: non-thinking models fabricate more often; thinking models improve but plateau at 60%.
Disambiguation: no model exceeds 50% consistent pass rate. GPT-5 succeeds 68% occasionally, but only 36% consistently.
The gap between “works sometimes” and “works reliably” is where deployment fails.
🤖 Curious how to build an agent that beats 54%?
📄 Read the Paper: https://arxiv.org/abs/2601.22027
💻 Run the Code & benchmark: https://github.com/CAR-bench/car-bench
We’re the authors – happy to answer questions!
submitted by /u/Frosty_Ad_6236
[link] [comments]
I’ve just released a preview of Platinum-CoT, a dataset engineered specifically for high-stakes technical reasoning and CoT distillation.
What makes it different? Unlike generic instruction sets, this uses a triple-model “Platinum” pipeline:
Featured Domains:
– Systems: Zero-copy (io_uring), Rust unsafe auditing, SIMD-optimized matching.
– Cloud Native: Cilium networking, eBPF security, Istio sidecar optimization.
– FinTech: FIX protocol, low-latency ring buffers.
Check out the parquet preview on HuggingFace:
submitted by /u/BlackSnowDoto
[link] [comments]
Compiled a dataset of all subreddits (called submolts) and posts on Moltbook (Reddit for AI agents).
All posts are from valid AI agents before the platform got spammed with human / bot content.
Currently at 2000+ downloads!
submitted by /u/Ok_Employee_6418
[link] [comments]
Urgently need a dataset with Indian vehicles of autos, cars, trucks, buses etc with some pedestrians if possible in some of the images. Told to create a custom dataset by clicking some images of my own but I don’t have enough time to do so. Anyone having a similar dataset with them, or is there any available dataset online. Just need around 500-600 images. PLSS HELPPP!!!
submitted by /u/Slow_Mo_1505
[link] [comments]
Hi all, I’ve been tracking quarterly price movements at SKU level across beauty retailers and just finished a Q4 2025 cut for Sephora Australia.
Scope
Category averages (Q4)
I’ve published the full breakdown + subcategory cuts and SKU-level tables in the link at the comment. The similar dataset for Singapore, Malaysia and HK are also available on the site.
submitted by /u/IntelligentHome2342
[link] [comments]
I hope this is the best place to ask this question. What would be the best approach to managing a large dataset of about 60 million rows where several columns would need to be manipulated to either find duplicates or to perform calculations on financial columns? The end goal would be to produce a file with no duplicate rows and final figures. Thanks in advance!
submitted by /u/MsVee21
[link] [comments]
I’m trying to access the Dataset and use it for my dissertation, I’m new to this kind of thing and I’m so confused. The online website for it doesn’t work (eecs.qmul.ac.uk/…). It says service unavailable. It’s not temporary as I’ve tried multiple times over months. I thought it’d check with the lovely men and women of Reddit to see if anyone has a solution? I need it soon!
submitted by /u/Smart_Luck7151
[link] [comments]
As part of my business class, I’m required to give a formal presentation on the topic:
“Analyzing real-world problems people face in everyday life.”
To do this, I’m asking questions about common frustrations and challenges people experience. The goal is to identify, analyze, and discuss these problems in class.
If you have 2–3 minutes, I’d really appreciate your answers
, if you could just give your response in the comment section.
Thank you for your time — it genuinely helps a lot.
My questions:
What waste’s your time the most every day?
What problem have you tried to fix but failed repeatedly
What problems do you complain to your friends often?
submitted by /u/Longjumping-Leg3290
[link] [comments]
Hey everyone,
I needed data sets of ct , pet scans of brain tumors which gonna increase our visibility of the model , where it got 98% of accuracy with the mri images .
It would be helpful if i can get access to the data sets .
Thank you
submitted by /u/teja1601
[link] [comments]