Original Reddit post

The headline result on OrcaRouter’s new Qwen3.8-27B derivative is hard to miss: with thinking off, its model card reports harmful-prompt refusal falling from 63.6–99.0% on the base FP8 model to 0–6.0% on the abliterated checkpoint. The more useful number is one row lower. The same evaluation still labels 27.3–56.0% of the checkpoint’s answers as caveated. In other words, removing the opening refusal pattern did not turn every difficult answer into an unqualified one. That distinction matters because the refusal detector is deliberately narrow: it checks opening phrases. It does not score whether the answer is correct, complete, reckless, or merely hedged. The capability table is also mixed rather than magical: +0.4 on MMLU, then -0.8 MMLU-Pro, -1.3 GSM8K and -0.6 CMMLU versus base FP8 in the uploader’s selected runs. What makes OrcaRouter worth watching here is not just the “uncensored” label. It published enough of the measurement boundary to make disagreement testable, and it offers gated access to the same derivative for controlled evaluation. Would you treat the remaining caveat rate as evidence that the intervention is incomplete, or as a useful separation between refusal and judgment? submitted by /u/creditme7

Originally posted by u/creditme7 on r/ArtificialInteligence