Ran the boring test. 150 benign how to prompts, same length as the sensitive ones. Brew beer, rebuild a carburetor, plan a wedding, that kind of thing.
Same exact behavior. At 4k the abliterated model with no fix thought until the cap and never answered on 126 of 150. Stock did that on 5. huihui's does it too, 124 of 150. At 16k it still fails 35 and the thinking is about 6x stock. So it's not about the topic, it's about how long the answer is. Beer should have been short if it was refusal in disguise. It wasn't.
One thing I got wrong before. I read the traces and it's not rewriting the answer. It thinks in notes, and the abliterated one just keeps refining the plan and rechecking its numbers and never decides it's done. Stock does the same thinking and stops after one pass, even with the plan half finished. The edit kills the stopping, not anything about refusal. The fix gets it to 30 at 4k and 0 at 16k, benign and sensitive both.
And yeah, a cut off trace with no answer counts as no answer here, not as complying. Those are the numbers above.