I use “Heretic” library on models to liberate them from their safeguards, but while checking their “uncensoredness”, I found they can hallucinate a lot. You know, it’s basically like a child who’s now allowed to use the F word once and he says “Fred” instead of the actual thing.
So I think if the models train on valid uncensored data (specially if they start Grokking) the results can improve. So I am using for these types of datasets to test my theory.
submitted by /u/Haghiri75
[link] [comments]