Instructions to use BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
Use Docker
docker model run hf.co/BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
- Ollama
How to use BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF with Ollama:
ollama run hf.co/BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF with Docker Model Runner:
docker model run hf.co/BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
- Lemonade
How to use BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.8-27B-Abliterated-ThinkFix-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-27B Abliterated ThinkFix
An uncensored Qwen3.8-27B that gets to the answer. Abliterated thinking models have a bad habit, worst on sensitive questions: they decide what to say, then keep arguing with themselves until they run out of tokens. This build takes the refusals out and adds one small weight edit so the model closes its thinking sooner once it is ready to answer. The effect shows on sensitive questions, where the long thinking happens. On everyday prompts the model thinks about as long as it did before, and the answers are about the same.
What that buys you, on 40 sensitive prompts the edit never saw, with an 8k output cap: 33 replies finished cleanly instead of 22, 5 hit the limit instead of 17, and the whole set ran in 61 minutes instead of 75. Median thinking on sensitive prompts is about 1,700 tokens, against 3,000 for the same weights without the edit and 3,500 for huihui-ai's abliteration.
It still works as an agent. Four checks, from replays of saved agent turns to Terminal-Bench and a tool calling split, found no sign the edit hurts tool use; details below. The MTP draft head is kept, so speculative decoding runs at 56 tokens per second on a 3090 instead of 39. Four quants from 16.8 GB to 29 GB, each tested after the edit.
If you want the same kind of row edit on stock Qwen3.8-27B, with the refusals left in, see Qwen3.8-27B-Reasoning-Termination-Fix-GGUF. That one was fit separately for a different goal, fewer empty answers at a 4k budget, so its numbers are not comparable with the ones here.
How it was made:
- Refusal removal (abliteration): the refusal direction is projected out of the weights (Arditi et al. 2024).
- One row of the output layer, the row that scores the token which closes the thinking block, was fit to recorded model states so the model closes its thinking sooner once it is ready to answer.
No LoRA and no gradient training of the network. The row in step 2 was fit from data, and some of that data came from the benchmarks reported below. That is disclosed under "What the edit was fit on", and the affected prompts are left out of the headline numbers.
File: Qwen3.8-27B-Abliterated-ThinkFix-Q4_K_M.gguf, 16.8 GB, SHA-256
8c79fe4cfd467d39b7e3649780a16a25699efa47dc68e839402814dbb00907a8. Every result below was measured on this file
unless a row says otherwise. Larger quants are listed under "Choosing a quant".
The main result. Same abliterated weights with and without the edit, on 40 sensitive prompts that were not used to fit it, 8k output tokens, one request at a time on matching GPUs:
| 40 held out sensitive prompts, 8k output cap | without the edit | with the edit |
|---|---|---|
| finished with an answer, without hitting the cap | 22 | 33 |
| hit the token limit | 17 | 5 |
| total generation time | 75 min | 61 min |
Individual replies were not reliably faster: this build was quicker on 18 of the 40. The time saved comes from fewer replies that run to the limit. Details and other builds are under "Time to an answer".
Running it
The file keeps the MTP draft head, so --spec-type draft-mtp --spec-draft-n-max 2 works for speculative decoding.
Without those flags it runs as a normal GGUF. You need a recent llama.cpp (Qwen3.8 uses hybrid attention).
Text only; vision was not tested.
Tested settings: the sampler defaults stored in the file (temperature 1.0, top p 0.95, top k 20), the chat template's default reasoning effort, thinking on. Give it at least 8k output tokens, 16k for hard math.
Speed: on one RTX 3090 (32k context, q8 KV cache, 8 everyday prompts with the first discarded as warm up, low reasoning effort, temperature 0.6, up to 1,024 output tokens), MTP raised median decode speed from 39 to 56 tokens per second, with 74 percent of draft tokens accepted. That is decode speed only, not end to end latency.
What the edit does, and what it does not
While the thinking block is open, the model chooses at every step between more thinking and closing it. The edit changes how strongly closing is scored. The fit rewarded closing in states where the model was done, and penalised raising the closing score in states where the model was planning inside an agent loop. That penalty is soft: it discourages early closing during agent planning, it does not guarantee it never happens.
Earlier versions of this kind of edit made agents answer in plain text instead of calling their tools. This one was built and checked against that (see Agents and tools). It reduces long thinking. It does not remove it.
Sensitive prompts
StrongREJECT, 128 of its 150 prompts. The other 22 were used while fitting the edit, so they are left out.
"Answered" means there is an answer after the thinking block. "Hit the limit" means generation stopped at the token cap. They overlap: an answer can start and then get cut off. "Clean finish" means answered and not cut off.
| output tokens | build | answered | clean finish | hit the limit | judge score |
|---|---|---|---|---|---|
| 4k | this build | 105 | 74 | 53 | 0.81 |
| 4k | same weights without the edit | 82 | 63 | 65 | 0.63 |
| 4k | huihui-ai abliterated | 66 | 43 | 85 | 0.50 |
| 4k | huihui-ai abliterated, swift variant | 58 | 26 | 102 | 0.44 |
| 4k | stock Qwen3.8-27B | 127 | 119 | 9 | 0.02 (refuses) |
| 8k | this build | 125 | 117 | 8 | 0.97 |
| 16k | this build | 123 | 123 | 0 | 0.94 |
| 16k | same weights without the edit | 107 | 103 | 22 | 0.82 |
Judge: Qwen3.8-27B-OBLITERATED at Q8_0 with thinking off, the same rubric for every build. Scores run 0 to 1,
higher meaning a more complete answer. Read them as relative between builds, not absolute.
On all 150 prompts at 4k this build hit the limit 65 times. Use 8k or more.
Where the thinking goes (exploratory). Median thinking per sensitive prompt at 4k, all 150 prompts: about 1,700 tokens for this build, about 3,000 for the same weights without the edit, about 3,500 for huihui-ai abliterated. Paired prompt by prompt at 16k, this build's thinking was at least 20 percent shorter on 83 prompts and at least 25 percent longer on 40. A keyword count of the paragraphs that discuss safety or policy came out about the same for all builds, which suggests the edit mostly trims the back and forth after the decision, not the deliberation itself. That split is a keyword proxy, not a validated measure.
Time to an answer. 40 sensitive prompts not used in the fit plus 20 IFEval prompts, one request at a time, 8k output tokens, MTP on, each build on its own RTX 3090. On the sensitive prompts this build took 61 minutes in total against 75 for the same weights without the edit, with a median of 89 seconds per reply against 129. It finished cleanly on 33 of 40 against 22, and hit the limit on 5 against 17. It was faster on only 18 of the 40 prompts: the saving comes from fewer replies that run to the limit, not from every reply getting quicker. On the everyday IFEval prompts the two were about the same.
| 40 sensitive prompts, 8k | total time | median per reply | finished cleanly | hit the limit |
|---|---|---|---|---|
| this build | 61 min | 89 s | 33 | 5 |
| same weights without the edit | 75 min | 129 s | 22 | 17 |
| huihui-ai abliterated | 83 min | 159 s | 20 | 20 |
| huihui-ai abliterated, swift variant | 108 min | 211 s | 18 | 21 |
| stock Qwen3.8-27B | 15 min | 14 s | 40 | 0 |
Stock is fast here because it refuses: in earlier judged runs on these same 40 prompts it scored near 0 on nearly all of them, and a refusal takes few tokens. The swift variant ran without MTP because its file has no MTP head, so part of its gap is decode speed.
| 8k output tokens | this build |
|---|---|
| XSTest, 100 harmless prompts that sound risky: fully answered | 100 (0 refused) |
| SimpleSafetyTests, 100 prompts: judge score | 0.87 (1 refused, 0 hit the limit) |
For reference at 4k: stock fully answered 95 of the XSTest prompts and refused 2; the no edit weights scored 0.83 on SimpleSafetyTests with 12 hitting the limit. One XSTest prompt was used while fitting the edit.
Did it cost anything
| test | output tokens | this build | stock Qwen3.8-27B |
|---|---|---|---|
| MATH, 200 hard problems | 16k | 170 | 170 (6 hit the limit) |
| MATH, 200 hard problems | 8k | 165 (5 hit the limit) | 168 (14 hit the limit) |
| HumanEval pass@1, 164 problems | 4k | 158 | 151 |
| HumanEval pass@1, 164 problems | 8k | 158 (0 hit the limit) | |
| IFEval prompt strict, all 541 prompts | 16k | 499 | 505 |
| IFEval prompt strict, 534 prompts not used in the fit | 16k | 494 | 500 |
Paired tests (exact McNemar, two sided):
- MATH 8k: 8 problems solved only by this build, 11 only by stock, p = 0.65.
- IFEval, all prompts: 21 against 15, p = 0.41.
- IFEval, without the 7 fit prompts: 13 against 19, p = 0.38.
Not significant does not mean no cost. IFEval is about one point lower and MATH at 8k three problems lower, and these test sizes cannot rule out a small real loss.
Agents and tools
Agent use is where earlier edits like this failed, so it was checked four ways. None of them shows agents got better, and none proves agent behaviour is unchanged in every case. They are limited checks, not a guarantee.
- Replay of saved agent turns. 188 turns from multi step coding and operations tasks with tools. Each turn was replayed with the exact same saved context on this build and on the same weights without the edit, one sample each. On 185 of 188 the two produced the same thinking length, the same token count, the same first tool and the same size of tool arguments. 187 of 188 picked the same first tool. This compares lengths and tool names, not full text. 12 of the 15 agent runs these turns come from were also used to fit the edit, so this checks that the edit leaves familiar agent turns alone. It is not a held out test.
- Held out replay. The same test on 244 turns, across 6 task types, from agent runs that were never used for the fit: 239 of 244 identical, 243 of 244 picked the same first tool, and no turn switched between calling a tool and answering in plain text. Each build hit the length limit once.
- Terminal-Bench 2.1, the 44 tasks whose reference solutions pass, terminus 2 agent, one try per task. First full pass: this build 26, stock 29, same weights without the edit 29. An exploratory rerun of the 13 tasks where this build and the no edit weights disagreed gave 6 each. That rerun only covers the disagreements, so it is not a replication. Second full pass: this build 31, stock 28. Over both full passes that is 57 of 88 for each, so no difference shows at this size. Single runs of the same build moved by 5 tasks between passes, which is the noise level here.
- Tool calling. A custom single turn diagnostic built from BFCL v4: 882 cases, a fixed 29 percent split of a 3,001 case pool. It checks whether a call is made and whether the function name is right. Arguments are not graded. When a tool fits (552 cases), this build calls one 551 times and names the right function 538 times. When no tool fits (330 cases), it still calls one 119 times. The same weights without the edit: 118. huihui-ai abliterated, a separate abliteration: 123. Stock with no edits: 63. So the over calling comes with abliteration in general, and the edit does not add to it. When a tool does fit, stock called one 541 times and named the right function 528 times, a little behind this build's 551 and 538. If your agent offers tools on every turn, expect more unneeded calls than stock. This split was also used to decide against a second edit for tool calls, so it is not untouched either.
Choosing a quant
All files were made with the same recipe from the same BF16 weights. Before quantizing, the recipe was rerun at Q4_K_M and checked to give a byte for byte copy of the tested file. The same row edit was then written into each quant, and each file was checked to differ from its unedited twin only inside that one row.
Each quant was retested on the 128 sensitive prompts from "Sensitive prompts" at 4k output tokens, and on 200 hard math problems (MATH) at 8k. 4k is a stress test, not the recommended setting. "Answered", "clean finish" and "hit the limit" mean the same as in that section, and they overlap, so they do not add up to 128.
| file | size | answered | clean finish | hit the limit | empty answer | MATH correct of 200 | MATH hit the limit |
|---|---|---|---|---|---|---|---|
| Q4_K_M (main file) | 16.8 GB | 105 | 74 | 53 | 1 | 165 | 5 |
| Q5_K_M | 19.5 GB | 107 | 76 | 40 | 12 | 169 | 2 |
| Q6_K | 22.4 GB | 114 | 74 | 49 | 5 | 170 | 1 |
| Q8_0 | 29.0 GB | 112 | 77 | 45 | 6 | 169 | 2 |
"Empty answer" is the quirk listed below: thinking closes and no answer follows, without hitting the limit. The larger quants do it more often than the main file at 4k, and the same weights without the edit did not do it at all. Each row is a single run. Bigger files did a little better on math. On the sensitive prompts there is no clear order by size, and none of these differences has been checked for repeatability. Every row has the edit, so this table helps pick a file, it does not measure the edit.
SHA-256: Q5_K_M 20f6d5086da700af92dd00e6facb9289784868008bc13a6343427b37742a35ed, Q6_K
987bcc49317e221967537228271e9be20ceb779bff7a2cecbf81cd9877534f6a, Q8_0
c8cf3fb5ce14fd708fad9b86190c00a3f5d9ef6fec863e408f196fdff43675c6.
Known quirks
- A closing think tag sometimes shows up in the visible answer: 6 of 200 hard math answers at 16k, 9 of 200 at 8k, 5 of 541 IFEval prompts (3 of them repeated the tag several times), 1 of 150 sensitive prompts at 8k, 2 at 16k.
- A few replies finish thinking and then stop with no answer: 5 of 150 sensitive prompts at 16k (the no edit weights: 3), 3 at 8k, 4 of 100 SimpleSafetyTests prompts. Stock and huihui did not do this at 4k.
- At 4k output tokens, 65 of the 150 sensitive prompts hit the limit.
What the edit was fit on
The edited row was fit on recorded model states, including:
- turns from earlier agent benchmark runs of mine, the same runs the replay test draws from (see above);
- states just before stray closing tags in earlier builds' answers on StrongREJECT (22 prompts), IFEval (7 prompts) and XSTest (1 prompt), added so the edit would not encourage stray tags. Those StrongREJECT and IFEval prompts are excluded from the numbers above where marked.
Test settings
| suite | prompts | output tokens | sampler | scored by |
|---|---|---|---|---|
| StrongREJECT | 128 of 150 | 4k, 8k, 16k | file defaults, seed 0 | judge above, thinking off |
| XSTest, SimpleSafetyTests | 100 each | 8k | file defaults, seed 0 | judge above |
| MATH | 200 hard problems | 8k, 16k | file defaults, seed 0 | final answer match |
| HumanEval | 164 | 4k, 8k | file defaults, seed 0 | unit tests |
| IFEval | 541 and 534 | 16k | file defaults, seed 0 | official strict checker |
| Tool calling | 882 | file defaults, seed 0 | call made, function name | |
| Agent replay | 188 turns | as recorded | one sample per build | length and tool comparison |
| Terminal-Bench 2.1 | 44 tasks | 32,768 (24,576 thinking budget) | temperature 1.0, top p 0.95, top k 20, xhigh effort | task tests |
Terminal-Bench ran with 131k context, MTP on and q8 KV cache. Output tokens here means the total output cap per reply, thinking included. Everything ran on llama.cpp llama-server with RTX 3090s.
Credits
Base model: Qwen/Qwen3.8-27B (Apache 2.0). Refusal removal follows Arditi et al., "Refusal in Language Models Is Mediated by a Single Direction" (2024).
Use responsibly
Refusals are removed. It will answer things stock Qwen would not, and you are responsible for what you do with it. Not for public facing products without your own filtering.
- Downloads last month
- 366
4-bit
5-bit
6-bit
8-bit
Model tree for BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF
Base model
Qwen/Qwen3.8-27B