99 Days Of AI Answer Engine Responses To A Fixed 16-question Set: 24,882 Scored Answers From OpenAI Search, Gemini And Claude, Plus A 10-model No-web Control (CC BY 4.0)

Disclosure: this is my own dataset. I built and run the measurement, and the subject of the measurement is my own pen name, so treat the topic with that in mind. The data itself is mechanical: same 16 questions, every day, three answer engines with web access, answers scored with a fixed rubric.

Dataset (Hugging Face, CC BY 4.0): https://huggingface.co/datasets/marintkael/ai-citation-fidelity

Five configs:

  • default: 24,882 scored answers including the no-web control channel
  • claude_web: 3,279 answers from a separate Claude web search panel
  • questions: the 16 questions with category and channel
  • data_gaps: register of measurement gaps (provider outages, method changes), because a gap is not a zero
  • daily_channels: daily time series split into direct, long tail and discovery channels

Collection window 2026-05-13 to 2026-08-19 (99 measurement days). The no-web control is 6,416 blind answers from 10 models without web access, useful for separating retrieval effects from training data. Scoring code and figure scripts: https://github.com/marintkael/marin-research-tools/tree/main/reports/03-found-not-recommended

Written report describing method and findings, if you want context before touching the parquet: https://doi.org/10.5281/zenodo.22015495

submitted by /u/marintkael
[link] [comments]

Leave a Reply

Your email address will not be published. Required fields are marked *