[Synthetic] [Self-Promotion] Research-Based CKD Dataset (200K Patients, 82 Clinical Features) For Machine Learning & Healthcare Analytics

Hi community,

I recently published a research-based synthetic Chronic Kidney Disease (CKD) dataset on Kaggle after spending several weeks studying clinical guidelines and epidemiological literature.

The motivation came from a common challenge I encountered: many publicly available CKD datasets contain only a few hundred patient records and a limited number of clinical variables, making them less suitable for building and evaluating modern machine learning models.

Dataset Highlights

• 200,000 synthetic patient records

• 82 clinically meaningful features

• Research-informed design using published clinical guidelines and epidemiological evidence

• Covers demographics, lifestyle, medical history, vital signs, kidney biomarkers, medications, frailty, healthcare utilization, and clinical outcomes

• Includes CKD stage, kidney failure risk, dialysis requirement, and hospitalization risk

• Machine learning and healthcare analytics ready

The dataset is completely synthetic and contains no real patient information. It was created for educational purposes, machine learning experiments, healthcare analytics, and research.

I’d really appreciate feedback from the community.

Some questions I’d love your thoughts on:

• Are there any important CKD-related variables you think are missing?

• What types of ML or analytics projects would you build with this dataset?

• What would you improve in a future version?

Kaggle Dataset:

https://www.kaggle.com/datasets/mohankrishnathalla/chronic-kidney-disease-risk-dataset-2026

Thanks for taking the time to check it out. I’m happy to answer questions about the design process or discuss future improvements.

submitted by /u/Mohan137
[link] [comments]

Leave a Reply

Your email address will not be published. Required fields are marked *