{"id":42370,"date":"2026-08-28T17:28:04","date_gmt":"2026-08-28T15:28:04","guid":{"rendered":"https:\/\/www.graviton.at\/letterswaplibrary\/im-stuck-finding-usable-historical-data-for-a-bayesian-pr-risk-model-looking-for-advice-on-how-to-proceed\/"},"modified":"2026-08-28T17:28:04","modified_gmt":"2026-08-28T15:28:04","slug":"im-stuck-finding-usable-historical-data-for-a-bayesian-pr-risk-model-looking-for-advice-on-how-to-proceed","status":"publish","type":"post","link":"https:\/\/www.graviton.at\/letterswaplibrary\/im-stuck-finding-usable-historical-data-for-a-bayesian-pr-risk-model-looking-for-advice-on-how-to-proceed\/","title":{"rendered":"I\u2019m Stuck Finding Usable Historical Data For A Bayesian PR Risk Model \u2014 Looking For Advice On How To Proceed"},"content":{"rendered":"<p><!-- SC_OFF --><\/p>\n<div class=\"md\">\n<p>Hi everyone,<\/p>\n<p>I\u2019m a student working on a research project on risk-aware GitHub PR review. I\u2019m doing the project mostly on my own and I don\u2019t have access to a research lab, large compute budget, or people who can manually annotate thousands of PRs, so I\u2019m trying to find a practical approach that I can actually finish.<\/p>\n<p>The idea is to take a GitHub PR and estimate four types of risk:<\/p>\n<ol>\n<li>\n<p>Bug \/ correctness<\/p>\n<\/li>\n<li>\n<p>Security<\/p>\n<\/li>\n<li>\n<p>Compatibility<\/p>\n<\/li>\n<li>\n<p>Cross-system \/ integration<\/p>\n<\/li>\n<\/ol>\n<p>The architecture I\u2019m working with has four separate risk models. They share the same PR characteristics\/features, but each risk model has its own historical data, prior, and evidence.<\/p>\n<p>My main problem is the historical data needed for those priors.<\/p>\n<p>At first, I looked for a single PR dataset where I could get reliable PR-level outcomes for all four risks. I couldn\u2019t find one.<\/p>\n<p>I then tried looking for separate datasets for each individual risk model. I thought this would solve the problem, but I keep finding datasets where the labels look relevant at first but don&#8217;t actually represent the outcome I need.<\/p>\n<p>For example, SEVRA-plus looked very promising for the Security model:<\/p>\n<p><a href=\"https:\/\/huggingface.co\/datasets\/RedAI4Code\/SEVRA-plus\">https:\/\/huggingface.co\/datasets\/RedAI4Code\/SEVRA-plus<\/a><\/p>\n<p>It contains security-related PR examples with vulnerability\/CWE information, but the malicious PRs are deliberately constructed by reversing real CVE security fixes. So although they are useful for evaluating or studying security vulnerabilities, I don&#8217;t think I can use their class distribution directly as a real-world prior for ordinary GitHub PRs.<\/p>\n<p>I\u2019ve run into similar issues with other datasets:<\/p>\n<p>&#8211; some label the linked issue rather than the PR implementation,<\/p>\n<p>&#8211; some label review comments rather than actual PR outcomes,<\/p>\n<p>&#8211; some contain artificially constructed vulnerable\/failing PRs,<\/p>\n<p>&#8211; some only give merge\/close status, which doesn\u2019t tell me whether the PR itself was buggy, vulnerable, incompatible, etc.<\/p>\n<p>The distinction between the issue and the PR is especially important for what I am trying to do.<\/p>\n<p>For example, imagine a maintainer opens a security issue, someone creates a PR to fix it, but the PR implementation itself contains a correctness bug and gets rejected. For my problem, I need to know the nature of the PR, not simply inherit the security label from the original issue.<\/p>\n<p>Similarly, a PR could be opened to fix a small bug, get merged, and then later cause a compatibility problem. Again, I care about what happened because of the PR implementation, not just why the PR was originally opened.<\/p>\n<p>Because I couldn&#8217;t find a dataset that directly gives me what I need, I tried a practical compromise.<\/p>\n<p>I took 96 real PRs from SWE-Review-Chat, filtered them for sufficient evidence, and used an LLM to annotate the four risk states from the information available in the PR record, such as the description, review discussion, diff context, tests, and lifecycle information.<\/p>\n<p>I\u2019m treating these as weak\/model-assisted labels rather than independent ground truth.<\/p>\n<p>The resulting usable outcomes are:<\/p>\n<p>Bug:<\/p>\n<p>30 present \/ 9 absent<\/p>\n<p>Security:<\/p>\n<p>1 present \/ 7 absent<\/p>\n<p>Compatibility:<\/p>\n<p>5 present \/ 12 absent<\/p>\n<p>Cross-system:<\/p>\n<p>4 present \/ 8 absent<\/p>\n<p>So now I feel like I\u2019ve hit a wall.<\/p>\n<p>I can keep searching for datasets, but so far I haven&#8217;t found anything that solves the underlying problem. I also don&#8217;t have the resources to manually establish reliable ground truth for thousands of PRs.<\/p>\n<p>I\u2019m therefore looking for advice on &#8220;how I should move forward from here&#8221;.<\/p>\n<p>Should I continue with the small real dataset I have and explicitly model the uncertainty caused by the sparse risks?<\/p>\n<p>Should I rely on LLM-assisted annotations of real PRs as a practical research compromise, or is there a better low-resource approach that I am missing?<\/p>\n<p>Or is there a completely different way of constructing the historical priors that would make more sense for this problem?<\/p>\n<p>I\u2019m not looking for a perfect dataset at this point. I\u2019m mainly looking for a practical and defensible way to move forward given that I\u2019m a student doing this alone with limited time and resources.<\/p>\n<p>If anyone has worked on GitHub PR datasets, Mining Software Repositories, empirical software engineering, code-review research, or Bayesian risk modelling, I would really appreciate any advice on what you would do in this situation.<\/p>\n<p>Thanks!<\/p>\n<\/div>\n<p><!-- SC_ON -->   submitted by   <a href=\"https:\/\/www.reddit.com\/user\/Accomplished-Fun4629\"> \/u\/Accomplished-Fun4629 <\/a> <br \/> <span><a href=\"https:\/\/www.reddit.com\/r\/datasets\/comments\/1w0tggn\/im_stuck_finding_usable_historical_data_for_a\/\">[link]<\/a><\/span>   <span><a href=\"https:\/\/www.reddit.com\/r\/datasets\/comments\/1w0tggn\/im_stuck_finding_usable_historical_data_for_a\/\">[comments]<\/a><\/span><\/p><div class='watch-action'><div class='watch-position align-right'><div class='action-like'><a class='lbg-style1 like-42370 jlk' href='javascript:void(0)' data-task='like' data-post_id='42370' data-nonce='7f483321c1' rel='nofollow'><img class='wti-pixel' src='https:\/\/www.graviton.at\/letterswaplibrary\/wp-content\/plugins\/wti-like-post\/images\/pixel.gif' title='Like' \/><span class='lc-42370 lc'>0<\/span><\/a><\/div><\/div> <div class='status-42370 status align-right'><\/div><\/div><div class='wti-clear'><\/div>","protected":false},"excerpt":{"rendered":"<p>Hi everyone, I\u2019m a student working on a research project on risk-aware GitHub PR review. I\u2019m doing&#8230;<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[85],"tags":[],"class_list":["post-42370","post","type-post","status-publish","format-standard","hentry","category-datatards","wpcat-85-id"],"_links":{"self":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts\/42370","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/comments?post=42370"}],"version-history":[{"count":0,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/posts\/42370\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/media?parent=42370"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/categories?post=42370"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.graviton.at\/letterswaplibrary\/wp-json\/wp\/v2\/tags?post=42370"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}