i everyone!
I built and curated a small, high-density synthetic dataset focused on academic research methodology in Spanish to help address the lack of native, non-translated academic datasets for LLM fine-tuning and alignment.
Key Specs
- Repository:
Januka2009/GPT5.6_SOL_INVESTIGACIONon Hugging Face. - Format & Size: 400 multi-turn conversations in JSONL/Parquet.
- Domain Balance: Stratified 90/10 train/test split across 4 core areas: Health, Social Sciences, STEM, and Economics.
- License: Creative Commons Attribution 4.0 (CC BY 4.0).
- Quality Check: Audited with an LLM-as-a-judge workflow (4.99/5 score) to strip conversational AI fluff and verify methodological accuracy.
It is ideal as seed data for LoRA/SFT experiments, academic tone alignment, or domain evaluation.
Any feedback, suggestions for new disciplines, or pull requests are greatly appreciated!
submitted by /u/Januka208338475
[link] [comments]