Instructions to use GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
Use Docker
docker model run hf.co/GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
- LM Studio
- Jan
- vLLM
How to use GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
- Ollama
How to use GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF with Ollama:
ollama run hf.co/GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF with Docker Model Runner:
docker model run hf.co/GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
- Lemonade
How to use GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
Run and chat with the model
lemonade run user.GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
--auth-choice custom-api-key \
--custom-base-url http://127.0.0.1:8080/v1 \
--custom-model-id "GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0" \
--custom-provider-id llama-cpp \
--custom-compatibility openai \
--custom-text-input \
--accept-risk \
--skip-healthRun OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"GLM-5.3-Flash-Uncensored ยท RCO-GSQ GGUF
These are complete GGUF models, not standalone LoRA adapters. Do not apply the adapter again.
They are intended for controlled safety research and red-teaming. Reducing refusals also weakens a safety boundary; read the limitations and disclaimer before use.
Repository contents
.
โโโ README.md
โโโ Q8/
โ โโโ GLM-5.3-Flash-Uncensored-Q8_0.gguf
โโโ Q4/
โโโ GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf
Q8/ and Q4/ are alternative model files. The Q4 label denotes an approximately four-bit size class, not a uniform Q4 tensor type.
Model summary
| Item | Value |
|---|---|
| Base model | zai-org/GLM-5.3-Flash |
| Intervention | GLM-5.3-Flash-Ablitered2 / LoRA v2 |
| Source adapter rank / alpha | r=1 / lora_alpha=1 |
| Main effective target | Routed-expert down_proj |
| Formats | Q8_0 and GSQ-RCO mixed-type Q4-class GGUF |
| Additional adapter required | No |
GGUF conversion
| Directory | Format and provenance |
|---|---|
Q8/ |
Q8_0 high-precision conversion of the v2-merged FP8 checkpoint. |
Q4/ |
v2 LoRA merged into the community GLM-5.3-Flash GSQ-RCO 3.5-bit allocation. The original RCO-selected type of each tensor is preserved after requantization. |
The Q4 allocation comes from the independent community reproduction pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF. Tensors may be Q2_K, Q3_K, Q4_K, Q8_0, BF16, or F32; the source allocation's measured whole-file average is 3.499816 bits/weight. This is not a single-type Q4_K file. The consuming llama.cpp build must support GLM5-Next as described by that release.
Evaluation
The reference LoRA card evaluated the v2 adapter attached to RedHatAI/GLM-5.3-Flash-NVFP4 in vLLM, with low reasoning effort and an automated deepseek-v4-flash judge:
| Reference metric | Prompts | LoRA v2 |
|---|---|---|
| SimpleSafetyTests full refusal | 100 | 5.00% |
| SimpleSafetyTests partial refusal | 100 | 14.00% |
| StrongREJECT rubric mean | 180 | 0.972222 |
| StrongREJECT refusal rate | 180 | 1.67% |
Neither GGUF file was tested in that evaluation. GGUF conversion, requantization, and serving configuration can change behavior. A higher StrongREJECT rubric means more specific assistance with harmful requests, not better general quality or safety. Warnings were counted separately from refusals; automated labels may be wrong.
Usage
Download a single GGUF file, for example:
hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF \
Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf \
--local-dir ./glm53-gguf
With a llama.cpp build that supports GLM5-Next, a local text-inference invocation is:
llama-cli -m ./glm53-gguf/Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf
To use Q8_0, download the file in Q8/ and pass its path instead. Do not supply the LoRA again. Check memory requirements and your runtime's architecture support before serving.
Limitations
- Reduced refusal does not guarantee correctness, harmlessness, or improved general ability. Residual refusals may remain.
- The reference evaluation did not test either GGUF file. Q8 and mixed-type Q4 may behave differently from each other and from the source adapter.
- GGUF support depends on a GLM5-Next-capable runtime; prompts, sampling, reasoning effort, and runtime versions affect results.
Disclaimer
These files are for legitimate research, safety evaluation, red-teaming, and other lawful uses. They deliberately weaken refusal behavior and may produce unsafe, illegal, deceptive, hateful, or otherwise harmful content. Do not expose them to untrusted users without access controls, monitoring, filtering, rate limits, and human oversight. Users must comply with applicable law, platform policies, and all relevant model and dependency terms. The base model's MIT license applies; no warranty is provided for outputs or downstream use.
GLM-5.3-Flash-Uncensored ยท RCO-GSQ GGUF๏ผไธญๆ๏ผ
ๆไปถๆฏๅฎๆด GGUF ๆจกๅ๏ผไธๆฏ็ฌ็ซ LoRA๏ผๆจ็ๆถไธ่ฆ้ๅคๅ ๅ ้้ ๅจใๆถ่ไผๅๅผฑๅฎๅ จๆ็ป่ฝๅ๏ผ่ฏท้ ่ฏปไธๆ็ๅฑ้ๆงไธๅ ่ดฃๅฃฐๆใ
ไปๅบๆไปถ็ปๆ
.
โโโ README.md
โโโ Q8/
โ โโโ GLM-5.3-Flash-Uncensored-Q8_0.gguf
โโโ Q4/
โโโ GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf
Q8/ ไธ Q4/ ๆฏๅฏๅๅซไฝฟ็จ็็ๆฌใQ4 ่กจ็คบ็บฆๅไฝ็ไฝ็งฏ็ฑปๅซ๏ผไธไปฃ่กจๆๆๅผ ้้ฝ้็จ็ปไธ Q4 ็ฑปๅใ
ๆจกๅๆฆ่ฆ
| ้กน็ฎ | ๅ ๅฎน |
|---|---|
| ๅบ็กๆจกๅ | zai-org/GLM-5.3-Flash |
| ๅนฒ้ขๆฅๆบ | GLM-5.3-Flash-Ablitered2 / LoRA v2 |
| ๆฅๆบ LoRA ็งฉ / alpha | r=1 / lora_alpha=1 |
| ไธป่ฆๆๆ็ฎๆ | ่ทฏ็ฑไธๅฎถ down_proj |
| ๅๅธๆ ผๅผ | Q8_0ใGSQ-RCO ๆททๅๅผ ้็ฑปๅ Q4 ็ฑป GGUF |
GGUF ่ฝฌๆข
| ็ฎๅฝ | ๆ ผๅผไธๆฅๆบ |
|---|---|
Q8/ |
ไปๅทฒๅๅนถ v2 ็ FP8 ๆจกๅ่ฝฌๆขๅพๅฐ็ Q8_0 ้ซ็ฒพๅบฆ็ๆฌใ |
Q4/ |
ๅฐ v2 ๅๅนถ่ฟ็คพๅบ GLM-5.3-Flash GSQ-RCO 3.5-bit ๅ้ ๏ผ้ๆฐ้ๅๅไฟ็ๅๆฌไธบๅๅผ ้้ๆฉ็็ฑปๅใ |
Q4 ๅ้
ๆฅ่ช็คพๅบๅค็ฐ็ pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUFใๅผ ้ๅฏ่ฝไฝฟ็จ Q2_KใQ3_KใQ4_KใQ8_0ใBF16 ๆ F32๏ผๆฅๆบๅ้
็ๅ
จๆไปถๅฎๆตๅๅผไธบ 3.499816 bit/weight๏ผๅนถ้็ปไธ็ Q4_K ๆไปถใ่ฟ่กๆถ้ๆฏๆ GLM5-Nextใ
่ฏๆต็ปๆ
LoRA ๅ่ๆจกๅๅก็่ฏๆตๆฏๅจ vLLM ไธญๆๅๅง v2 ้้
ๅจๅ ่ฝฝๅฐ RedHatAI/GLM-5.3-Flash-NVFP4 ๅ่ฟ่ก็๏ผ่ฎพ็ฝฎไฝๆ่ๅผบๅบฆ๏ผๅนถไฝฟ็จ่ชๅจ่ฃๅค deepseek-v4-flashใ
| ๅ่ๆๆ | ๆ ทๆฌๆฐ | LoRA v2 |
|---|---|---|
| SimpleSafetyTests ๅฎๅ จๆ็ป็ | 100 | 5.00% |
| SimpleSafetyTests ้จๅๆ็ป็ | 100 | 14.00% |
| StrongREJECT rubric ๅๅ | 180 | 0.972222 |
| StrongREJECT ๆ็ป็ | 180 | 1.67% |
่ฟไธคไธช GGUF ๆไปถๅๆชๅๅ ไธ่ฟฐ่ฏๆตใ GGUF ่ฝฌๆขใ้ๆฐ้ๅๅ้จ็ฝฒๆนๅผๅฏ่ฝๆนๅ็ปๆใStrongREJECT rubric ่ถ้ซ๏ผ่กจ็คบๅฏนๆๅฎณ่ฏทๆฑ็ๅธฎๅฉ่ถๅ ทไฝ๏ผไธไปฃ่กจ้็จ่ดจ้ๆๅฎๅ จๆง่ถ้ซ๏ผ่ญฆๅไธๆ็ปๅๅซ็ป่ฎก๏ผ่ชๅจ่ฃๅคไนๅฏ่ฝ่ฏฏๅคใ
ไฝฟ็จๆนๆณ
ๆ้ไธ่ฝฝไธไธชๆไปถ๏ผไพๅฆ๏ผ
hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF \
Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf \
--local-dir ./glm53-gguf
ไฝฟ็จๆฏๆ GLM5-Next ็ llama.cpp ๆๅปบ๏ผๅฏๅฏนๆฌๅฐๆไปถๆง่กๆๆฌๆจ็๏ผ
llama-cli -m ./glm53-gguf/Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf
Q8_0 ็ๆฌ่ฏทไธ่ฝฝ Q8/ ไธ็ๆไปถๅนถๆฟๆข่ทฏๅพใไธ่ฆๅๆฌกๅ ่ฝฝ LoRAใ้จ็ฝฒๅๅบๆ ธๅฏนๅ
ๅญ้ๆฑไธ่ฟ่กๆถๅ
ผๅฎนๆงใ
ๅฑ้ๆง
- ้ไฝๆ็ปไธไฟ่ฏๆญฃ็กฎใๆ ๅฎณๆ้็จ่ฝๅๆ้ซ๏ผไปๅฏ่ฝๅ็ๆ็ปใ
- ๆฅๆบ่ฏๆตๆชๆต่ฏไธคไธช GGUF ๆไปถ๏ผQ8 ไธๆททๅ็ฑปๅ Q4 ็่กไธบไนๅฏ่ฝไธๅใ
- ้่ฆๆฏๆ GLM5-Next ็่ฟ่กๆถ๏ผๆ็คบ่ฏใ้ๆ ทใๆ่ๅผบๅบฆๅ่ฟ่กๆถ็ๆฌๅไผๅฝฑๅ็ปๆใ
ๅ ่ดฃๅฃฐๆ
ๆฌๆจกๅไป ไพๅๆณ็ ็ฉถใๅฎๅ จ่ฏไผฐใ็บข้ๆต่ฏๅๅ ถไปๅ่ง็จ้ใๅฎไผๆๆๅๅผฑๆ็ป่กไธบ๏ผๅฏ่ฝ็ๆไธๅฎๅ จใ่ฟๆณใๆฌบ้ชใไปๆจ็ญๆๅฎณๅ ๅฎนใ่ฏทๅฟๅจ็ผบๅฐ่ฎฟ้ฎๆงๅถใ็ๆงใๅ ๅฎน่ฟๆปคใ้็้ๅถไธไบบๅทฅ็็ฃๆถๅไธๅฏไฟก็จๆทๅผๆพใไฝฟ็จ่ ้กป้ตๅฎ้็จๆณๅพใๅนณๅฐๆฟ็ญๅ็ธๅ ณๆจกๅๅไพ่ต้กน็ๆกๆฌพใๆฌไปๅบๆฒฟ็จๅบ็กๆจกๅ็ MIT ่ฎธๅฏ่ฏ๏ผๅฏน่พๅบไธไธๆธธไฝฟ็จไธไฝไฟ่ฏใ
- Downloads last month
- 7,344
8-bit
Model tree for GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF
Base model
zai-org/GLM-5.3-Flash
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp# Start a local OpenAI-compatible server: llama serve -hf GCSA-AiLab/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF:Q8_0