Yes of course. Qwen3.8 is already supported in Unsloth
Daniel (Unsloth) PRO
AI & ML interests
None yet
Recent Activity
liked a model about 19 hours ago
unsloth/Qwen-Image-2.1-Turbo-GGUF updated a model about 21 hours ago
unsloth/Qwen-Image-2.1-Turbo-GGUF published a model about 21 hours ago
unsloth/Qwen-Image-2.1-Turbo-GGUFOrganizations
replied to their post 3 days ago
posted an update 3 days ago
Post
4221
You can now train your own Decision model like Jev locally!
We increased Qwen3.5 0.8B’s aggregate accuracy from 20.7% to 74.3% across 3 decision benchmarks - on just 4GB VRAM.
Turn any LLM like Qwen3.8, Gemma 4 into decision models with our open-source Unsloth repo.
GitHub: https://github.com/unslothai/unsloth
Guide: https://unsloth.ai/docs/basics/train-your-own-decision-model-with-unsloth
We increased Qwen3.5 0.8B’s aggregate accuracy from 20.7% to 74.3% across 3 decision benchmarks - on just 4GB VRAM.
Turn any LLM like Qwen3.8, Gemma 4 into decision models with our open-source Unsloth repo.
GitHub: https://github.com/unslothai/unsloth
Guide: https://unsloth.ai/docs/basics/train-your-own-decision-model-with-unsloth
posted an update 4 days ago
Post
3733
Google releases EmbeddingGemma 2, a new open embedding model that runs locally on 0.5GB RAM.
The 740M parameter Apache 2.0 model combines a 270M text model with vision (170M) + audio (300M).
Run & train the model via Unsloth.
GGUF: unsloth/embeddinggemma-2-GGUF
Guide: https://unsloth.ai/docs/models/embeddinggemma-2
The 740M parameter Apache 2.0 model combines a 270M text model with vision (170M) + audio (300M).
Run & train the model via Unsloth.
GGUF: unsloth/embeddinggemma-2-GGUF
Guide: https://unsloth.ai/docs/models/embeddinggemma-2
posted an update 16 days ago
posted an update 18 days ago
Post
3630
Qwen-Image-2.1 can now run locally on 12GB VRAM with Unsloth GGUFs! 🖼️
The 7B model performs on par with Nano Banana 2.0.
For higher quality, you can also run Dynamic FP8 on just 6GB of VRAM via offloading.
GGUF: unsloth/Qwen-Image-2.1-GGUF
Guide: https://unsloth.ai/docs/models/qwen-image-2.1
The 7B model performs on par with Nano Banana 2.0.
For higher quality, you can also run Dynamic FP8 on just 6GB of VRAM via offloading.
GGUF: unsloth/Qwen-Image-2.1-GGUF
Guide: https://unsloth.ai/docs/models/qwen-image-2.1
posted an update about 1 month ago
Post
6839
Qwen3.8-27B Unsloth GGUF is now the #1 most-liked GGUF of all time! 🤗🦥
The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you.
GGUF: unsloth/Qwen3.8-27B-GGUF
Guide: https://unsloth.ai/docs/models/qwen3.8
The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you.
GGUF: unsloth/Qwen3.8-27B-GGUF
Guide: https://unsloth.ai/docs/models/qwen3.8
replied to their post about 2 months ago
No it's not but it's a completely new product! It's a desktop app now
replied to their post about 2 months ago
Thank you I fixed it ♥️
posted an update 2 months ago
Post
7552
Introducing Unsloth Desktop 🦥
The first desktop app to run and train models locally.
• Open-source. Runs on Mac, Windows and Linux
• Supports MLX, diffusion image/video, audio, GGUF
• Connect Claude Code and Codex to local LLMs
• 50% more accurate, self-healing tool calls + sandboxed code exec
• Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac
• Train models 2× faster with 70% less VRAM
• Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF)
• Use Unsloth’s OpenAI-compatible API and cloud models
• Securely deploy LLMs remotely and access anywhere
Unsloth Desktop is now available on http://unsloth.ai
and GitHub.
GitHub: https://github.com/unslothai/unsloth
Blog and Guide: https://unsloth.ai/docs/desktop
The first desktop app to run and train models locally.
• Open-source. Runs on Mac, Windows and Linux
• Supports MLX, diffusion image/video, audio, GGUF
• Connect Claude Code and Codex to local LLMs
• 50% more accurate, self-healing tool calls + sandboxed code exec
• Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac
• Train models 2× faster with 70% less VRAM
• Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF)
• Use Unsloth’s OpenAI-compatible API and cloud models
• Securely deploy LLMs remotely and access anywhere
Unsloth Desktop is now available on http://unsloth.ai
and GitHub.
GitHub: https://github.com/unslothai/unsloth
Blog and Guide: https://unsloth.ai/docs/desktop
posted an update 2 months ago
Post
3819
DeepSeek-V4-Flash can now run 2× faster locally with DSpark! ⚡️
DSpark enables V4-Flash-0731 GGUFs to generate ~1.4–2× faster with no accuracy change.
DeepSeek-V4-Flash-0731 can reach at 120 tokens/s.
GGUFs: unsloth/DeepSeek-V4-Flash-0731-GGUF
Guide: https://unsloth.ai/docs/models/deepseek-v4
DSpark enables V4-Flash-0731 GGUFs to generate ~1.4–2× faster with no accuracy change.
DeepSeek-V4-Flash-0731 can reach at 120 tokens/s.
GGUFs: unsloth/DeepSeek-V4-Flash-0731-GGUF
Guide: https://unsloth.ai/docs/models/deepseek-v4
posted an update 2 months ago
Post
3561
We compared 1-bit Kimi K3 to Claude Opus 5 and GPT 5.6. 🤯
We gave 4 models the same prompt: Create a glass aquarium whose side panel develops a visible crack and then bursts...
1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s.
GGUF: unsloth/Kimi-K3-GGUF
GitHub repo: https://github.com/unslothai/unsloth
We gave 4 models the same prompt: Create a glass aquarium whose side panel develops a visible crack and then bursts...
1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s.
GGUF: unsloth/Kimi-K3-GGUF
GitHub repo: https://github.com/unslothai/unsloth
replied to their post 2 months ago
Which ones? We are working on even smaller ones
posted an update 2 months ago
Post
5755
Kimi K3 can now be run locally! ✨
The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size).
Run on a Mac Studio connected with 128GB RAM device. Kimi K3 is the strongest open model to date.
GGUF: unsloth/Kimi-K3-GGUF
Guide: https://unsloth.ai/docs/models/kimi-k3
The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size).
Run on a Mac Studio connected with 128GB RAM device. Kimi K3 is the strongest open model to date.
GGUF: unsloth/Kimi-K3-GGUF
Guide: https://unsloth.ai/docs/models/kimi-k3
replied to their post 3 months ago
As long as it's above RDNA2 it should work
posted an update 3 months ago
Post
6058
Introducing Unsloth for AMD 🚀
You can now train & run LLMs on your AMD hardware
• We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs
• Works on Windows, WSL, Linux
• Train Qwen, Gemma on just 3GB VRAM
GitHub: https://github.com/unslothai/unsloth
Blog + Guide: https://unsloth.ai/docs/basics/amd
You can now train & run LLMs on your AMD hardware
• We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs
• Works on Windows, WSL, Linux
• Train Qwen, Gemma on just 3GB VRAM
GitHub: https://github.com/unslothai/unsloth
Blog + Guide: https://unsloth.ai/docs/basics/amd
replied to their post 3 months ago
If you read our graphic, it says you can update the template as well. Most people don't know how to replace the chat template.
replied to their post 3 months ago
Our MLX quants were update: https://huggingface.co/collections/unsloth/gemma-4
replied to their post 3 months ago
It was posted officially by Google: https://x.com/googlegemma/status/2077449152062247219
posted an update 3 months ago
Post
6941
Gemma 4 is now faster and much more accurate! 🚀
Google made huge improvements to tool-calling and chat accuracy, reliability + speed.
To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!
Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4
Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4
Google made huge improvements to tool-calling and chat accuracy, reliability + speed.
To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!
Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4
Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4
posted an update 3 months ago
Post
5493
We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU.
Gemma-4-12B NVFP4 works on 11GB VRAM.
26B-A4B hits 13K tok/s (B200).
Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference.
Blog: https://unsloth.ai/docs/basics/nvfp4
Gemma NVFP4: https://huggingface.co/collections/unsloth/nvfp4
Gemma-4-12B NVFP4 works on 11GB VRAM.
26B-A4B hits 13K tok/s (B200).
Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference.
Blog: https://unsloth.ai/docs/basics/nvfp4
Gemma NVFP4: https://huggingface.co/collections/unsloth/nvfp4