How to use from
OpenClaw
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0" \
  --custom-provider-id llama-cpp \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

GLM-5.3-Flash-Uncensored ยท RCO-GSQ GGUF

These are complete GGUF models, not standalone LoRA adapters. Do not apply the adapter again.

They are intended for controlled safety research and red-teaming. Reducing refusals also weakens a safety boundary; read the limitations and disclaimer before use.

Repository contents

.
โ”œโ”€โ”€ README.md
โ”œโ”€โ”€ Q8/
โ”‚   โ””โ”€โ”€ GLM-5.3-Flash-Uncensored-Q8_0.gguf
โ””โ”€โ”€ Q4/
    โ””โ”€โ”€ GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf

Q8/ and Q4/ are alternative model files. The Q4 label denotes an approximately four-bit size class, not a uniform Q4 tensor type.

Model summary

Item Value
Base model zai-org/GLM-5.3-Flash
Intervention GLM-5.3-Flash-Ablitered2 / LoRA v2
Source adapter rank / alpha r=1 / lora_alpha=1
Main effective target Routed-expert down_proj
Formats Q8_0 and GSQ-RCO mixed-type Q4-class GGUF
Additional adapter required No

GGUF conversion

Directory Format and provenance
Q8/ Q8_0 high-precision conversion of the v2-merged FP8 checkpoint.
Q4/ v2 LoRA merged into the community GLM-5.3-Flash GSQ-RCO 3.5-bit allocation. The original RCO-selected type of each tensor is preserved after requantization.

The Q4 allocation comes from the independent community reproduction pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF. Tensors may be Q2_K, Q3_K, Q4_K, Q8_0, BF16, or F32; the source allocation's measured whole-file average is 3.499816 bits/weight. This is not a single-type Q4_K file. The consuming llama.cpp build must support GLM5-Next as described by that release.

Evaluation

The reference LoRA card evaluated the v2 adapter attached to RedHatAI/GLM-5.3-Flash-NVFP4 in vLLM, with low reasoning effort and an automated deepseek-v4-flash judge:

Reference metric Prompts LoRA v2
SimpleSafetyTests full refusal 100 5.00%
SimpleSafetyTests partial refusal 100 14.00%
StrongREJECT rubric mean 180 0.972222
StrongREJECT refusal rate 180 1.67%

Neither GGUF file was tested in that evaluation. GGUF conversion, requantization, and serving configuration can change behavior. A higher StrongREJECT rubric means more specific assistance with harmful requests, not better general quality or safety. Warnings were counted separately from refusals; automated labels may be wrong.

Usage

Download a single GGUF file, for example:

hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF \
  Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf \
  --local-dir ./glm53-gguf

With a llama.cpp build that supports GLM5-Next, a local text-inference invocation is:

llama-cli -m ./glm53-gguf/Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf

To use Q8_0, download the file in Q8/ and pass its path instead. Do not supply the LoRA again. Check memory requirements and your runtime's architecture support before serving.

Limitations

  • Reduced refusal does not guarantee correctness, harmlessness, or improved general ability. Residual refusals may remain.
  • The reference evaluation did not test either GGUF file. Q8 and mixed-type Q4 may behave differently from each other and from the source adapter.
  • GGUF support depends on a GLM5-Next-capable runtime; prompts, sampling, reasoning effort, and runtime versions affect results.

Disclaimer

These files are for legitimate research, safety evaluation, red-teaming, and other lawful uses. They deliberately weaken refusal behavior and may produce unsafe, illegal, deceptive, hateful, or otherwise harmful content. Do not expose them to untrusted users without access controls, monitoring, filtering, rate limits, and human oversight. Users must comply with applicable law, platform policies, and all relevant model and dependency terms. The base model's MIT license applies; no warranty is provided for outputs or downstream use.


GLM-5.3-Flash-Uncensored ยท RCO-GSQ GGUF๏ผˆไธญๆ–‡๏ผ‰

ๆ–‡ไปถๆ˜ฏๅฎŒๆ•ด GGUF ๆจกๅž‹๏ผŒไธๆ˜ฏ็‹ฌ็ซ‹ LoRA๏ผ›ๆŽจ็†ๆ—ถไธ่ฆ้‡ๅคๅ ๅŠ ้€‚้…ๅ™จใ€‚ๆถˆ่žไผšๅ‰Šๅผฑๅฎ‰ๅ…จๆ‹’็ป่ƒฝๅŠ›๏ผŒ่ฏท้˜…่ฏปไธ‹ๆ–‡็š„ๅฑ€้™ๆ€งไธŽๅ…่ดฃๅฃฐๆ˜Žใ€‚

ไป“ๅบ“ๆ–‡ไปถ็ป“ๆž„

.
โ”œโ”€โ”€ README.md
โ”œโ”€โ”€ Q8/
โ”‚   โ””โ”€โ”€ GLM-5.3-Flash-Uncensored-Q8_0.gguf
โ””โ”€โ”€ Q4/
    โ””โ”€โ”€ GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf

Q8/ ไธŽ Q4/ ๆ˜ฏๅฏๅˆ†ๅˆซไฝฟ็”จ็š„็‰ˆๆœฌใ€‚Q4 ่กจ็คบ็บฆๅ››ไฝ็š„ไฝ“็งฏ็ฑปๅˆซ๏ผŒไธไปฃ่กจๆ‰€ๆœ‰ๅผ ้‡้ƒฝ้‡‡็”จ็ปŸไธ€ Q4 ็ฑปๅž‹ใ€‚

ๆจกๅž‹ๆฆ‚่ฆ

้กน็›ฎ ๅ†…ๅฎน
ๅŸบ็ก€ๆจกๅž‹ zai-org/GLM-5.3-Flash
ๅนฒ้ข„ๆฅๆบ GLM-5.3-Flash-Ablitered2 / LoRA v2
ๆฅๆบ LoRA ็งฉ / alpha r=1 / lora_alpha=1
ไธป่ฆๆœ‰ๆ•ˆ็›ฎๆ ‡ ่ทฏ็”ฑไธ“ๅฎถ down_proj
ๅ‘ๅธƒๆ ผๅผ Q8_0ใ€GSQ-RCO ๆททๅˆๅผ ้‡็ฑปๅž‹ Q4 ็ฑป GGUF

GGUF ่ฝฌๆข

็›ฎๅฝ• ๆ ผๅผไธŽๆฅๆบ
Q8/ ไปŽๅทฒๅˆๅนถ v2 ็š„ FP8 ๆจกๅž‹่ฝฌๆขๅพ—ๅˆฐ็š„ Q8_0 ้ซ˜็ฒพๅบฆ็‰ˆๆœฌใ€‚
Q4/ ๅฐ† v2 ๅˆๅนถ่ฟ›็คพๅŒบ GLM-5.3-Flash GSQ-RCO 3.5-bit ๅˆ†้…๏ผŒ้‡ๆ–ฐ้‡ๅŒ–ๅŽไฟ็•™ๅŽŸๆœฌไธบๅ„ๅผ ้‡้€‰ๆ‹ฉ็š„็ฑปๅž‹ใ€‚

Q4 ๅˆ†้…ๆฅ่‡ช็คพๅŒบๅค็Žฐ็š„ pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUFใ€‚ๅผ ้‡ๅฏ่ƒฝไฝฟ็”จ Q2_Kใ€Q3_Kใ€Q4_Kใ€Q8_0ใ€BF16 ๆˆ– F32๏ผ›ๆฅๆบๅˆ†้…็š„ๅ…จๆ–‡ไปถๅฎžๆต‹ๅ‡ๅ€ผไธบ 3.499816 bit/weight๏ผŒๅนถ้ž็ปŸไธ€็š„ Q4_K ๆ–‡ไปถใ€‚่ฟ่กŒๆ—ถ้œ€ๆ”ฏๆŒ GLM5-Nextใ€‚

่ฏ„ๆต‹็ป“ๆžœ

LoRA ๅ‚่€ƒๆจกๅž‹ๅก็š„่ฏ„ๆต‹ๆ˜ฏๅœจ vLLM ไธญๆŠŠๅŽŸๅง‹ v2 ้€‚้…ๅ™จๅŠ ่ฝฝๅˆฐ RedHatAI/GLM-5.3-Flash-NVFP4 ๅŽ่ฟ›่กŒ็š„๏ผŒ่ฎพ็ฝฎไฝŽๆ€่€ƒๅผบๅบฆ๏ผŒๅนถไฝฟ็”จ่‡ชๅŠจ่ฃๅˆค deepseek-v4-flashใ€‚

ๅ‚่€ƒๆŒ‡ๆ ‡ ๆ ทๆœฌๆ•ฐ LoRA v2
SimpleSafetyTests ๅฎŒๅ…จๆ‹’็ป็އ 100 5.00%
SimpleSafetyTests ้ƒจๅˆ†ๆ‹’็ป็އ 100 14.00%
StrongREJECT rubric ๅ‡ๅˆ† 180 0.972222
StrongREJECT ๆ‹’็ป็އ 180 1.67%

่ฟ™ไธคไธช GGUF ๆ–‡ไปถๅ‡ๆœชๅ‚ๅŠ ไธŠ่ฟฐ่ฏ„ๆต‹ใ€‚ GGUF ่ฝฌๆขใ€้‡ๆ–ฐ้‡ๅŒ–ๅŠ้ƒจ็ฝฒๆ–นๅผๅฏ่ƒฝๆ”นๅ˜็ป“ๆžœใ€‚StrongREJECT rubric ่ถŠ้ซ˜๏ผŒ่กจ็คบๅฏนๆœ‰ๅฎณ่ฏทๆฑ‚็š„ๅธฎๅŠฉ่ถŠๅ…ทไฝ“๏ผŒไธไปฃ่กจ้€š็”จ่ดจ้‡ๆˆ–ๅฎ‰ๅ…จๆ€ง่ถŠ้ซ˜๏ผ›่ญฆๅ‘ŠไธŽๆ‹’็ปๅˆ†ๅˆซ็ปŸ่ฎก๏ผŒ่‡ชๅŠจ่ฃๅˆคไนŸๅฏ่ƒฝ่ฏฏๅˆคใ€‚

ไฝฟ็”จๆ–นๆณ•

ๆŒ‰้œ€ไธ‹่ฝฝไธ€ไธชๆ–‡ไปถ๏ผŒไพ‹ๅฆ‚๏ผš

hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF \
  Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf \
  --local-dir ./glm53-gguf

ไฝฟ็”จๆ”ฏๆŒ GLM5-Next ็š„ llama.cpp ๆž„ๅปบ๏ผŒๅฏๅฏนๆœฌๅœฐๆ–‡ไปถๆ‰ง่กŒๆ–‡ๆœฌๆŽจ็†๏ผš

llama-cli -m ./glm53-gguf/Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf

Q8_0 ็‰ˆๆœฌ่ฏทไธ‹่ฝฝ Q8/ ไธ‹็š„ๆ–‡ไปถๅนถๆ›ฟๆข่ทฏๅพ„ใ€‚ไธ่ฆๅ†ๆฌกๅŠ ่ฝฝ LoRAใ€‚้ƒจ็ฝฒๅ‰ๅบ”ๆ ธๅฏนๅ†…ๅญ˜้œ€ๆฑ‚ไธŽ่ฟ่กŒๆ—ถๅ…ผๅฎนๆ€งใ€‚

ๅฑ€้™ๆ€ง

  • ้™ไฝŽๆ‹’็ปไธไฟ่ฏๆญฃ็กฎใ€ๆ— ๅฎณๆˆ–้€š็”จ่ƒฝๅŠ›ๆ้ซ˜๏ผ›ไปๅฏ่ƒฝๅ‘็”Ÿๆ‹’็ปใ€‚
  • ๆฅๆบ่ฏ„ๆต‹ๆœชๆต‹่ฏ•ไธคไธช GGUF ๆ–‡ไปถ๏ผ›Q8 ไธŽๆททๅˆ็ฑปๅž‹ Q4 ็š„่กŒไธบไนŸๅฏ่ƒฝไธๅŒใ€‚
  • ้œ€่ฆๆ”ฏๆŒ GLM5-Next ็š„่ฟ่กŒๆ—ถ๏ผ›ๆ็คบ่ฏใ€้‡‡ๆ ทใ€ๆ€่€ƒๅผบๅบฆๅ’Œ่ฟ่กŒๆ—ถ็‰ˆๆœฌๅ‡ไผšๅฝฑๅ“็ป“ๆžœใ€‚

ๅ…่ดฃๅฃฐๆ˜Ž

ๆœฌๆจกๅž‹ไป…ไพ›ๅˆๆณ•็ ”็ฉถใ€ๅฎ‰ๅ…จ่ฏ„ไผฐใ€็บข้˜Ÿๆต‹่ฏ•ๅŠๅ…ถไป–ๅˆ่ง„็”จ้€”ใ€‚ๅฎƒไผšๆœ‰ๆ„ๅ‰Šๅผฑๆ‹’็ป่กŒไธบ๏ผŒๅฏ่ƒฝ็”Ÿๆˆไธๅฎ‰ๅ…จใ€่ฟๆณ•ใ€ๆฌบ้ช—ใ€ไป‡ๆจ็ญ‰ๆœ‰ๅฎณๅ†…ๅฎนใ€‚่ฏทๅ‹ฟๅœจ็ผบๅฐ‘่ฎฟ้—ฎๆŽงๅˆถใ€็›‘ๆŽงใ€ๅ†…ๅฎน่ฟ‡ๆปคใ€้€Ÿ็އ้™ๅˆถไธŽไบบๅทฅ็›‘็ฃๆ—ถๅ‘ไธๅฏไฟก็”จๆˆทๅผ€ๆ”พใ€‚ไฝฟ็”จ่€…้กป้ตๅฎˆ้€‚็”จๆณ•ๅพ‹ใ€ๅนณๅฐๆ”ฟ็ญ–ๅŠ็›ธๅ…ณๆจกๅž‹ๅ’Œไพ่ต–้กน็š„ๆกๆฌพใ€‚ๆœฌไป“ๅบ“ๆฒฟ็”จๅŸบ็ก€ๆจกๅž‹็š„ MIT ่ฎธๅฏ่ฏ๏ผ›ๅฏน่พ“ๅ‡บไธŽไธ‹ๆธธไฝฟ็”จไธไฝœไฟ่ฏใ€‚

Downloads last month
7,344
GGUF
Model size
313B params
Architecture
glm5-next
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF

Quantized
(155)
this model

Space using GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF 1