[Data Release] Spanish Academic Research Methodology Dataset (400 Multiturn Records, CC BY 4.0)

i everyone!

I built and curated a small, high-density synthetic dataset focused on academic research methodology in Spanish to help address the lack of native, non-translated academic datasets for LLM fine-tuning and alignment.

Key Specs

  • Repository: Januka2009/GPT5.6_SOL_INVESTIGACION on Hugging Face.
  • Format & Size: 400 multi-turn conversations in JSONL/Parquet.
  • Domain Balance: Stratified 90/10 train/test split across 4 core areas: Health, Social Sciences, STEM, and Economics.
  • License: Creative Commons Attribution 4.0 (CC BY 4.0).
  • Quality Check: Audited with an LLM-as-a-judge workflow (4.99/5 score) to strip conversational AI fluff and verify methodological accuracy.

It is ideal as seed data for LoRA/SFT experiments, academic tone alignment, or domain evaluation.

Any feedback, suggestions for new disciplines, or pull requests are greatly appreciated!

submitted by /u/Januka208338475
[link] [comments]

Leave a Reply

Your email address will not be published. Required fields are marked *