Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up

All HF Hub posts

tardellirsΒ 
posted an update 2 days ago
view post
Post
5862
I built Model Pulse: daily download history for every model on the Hub πŸ“ˆ

A model page tells you its downloads for the last 30 days, but not how it got there. Model Pulse rebuilds the day-by-day history for 1.6M models, back to July 2024, from the daily snapshots in @cfahlgren1 's hub-stats dataset.

A few things the data shows:
- The Hub now serves ~100M model downloads a day, up from ~62M a year ago
- Vision-language models grew about 7x in that year, to 10.7M downloads a day
- 37% of Qwen3-8B's monthly downloads go to its 6,283 quantizations, fine-tunes and merges, not to the original

What you can do with it:
- Open any model, e.g. tardellirs/model-pulse
- Compare up to five models on one chart
- Browse weekly rankings: fastest growing, new this month, biggest families, top organizations
- Add a live badge to your model card (monthly downloads, sparkline and weekly trend)

The data is open too: modelpulse/model-pulse-data

I'd love feedback, especially from model authors: what would you want to see about your own models?
  • 7 replies
Β·
SeaWolf-AIΒ 
posted an update 3 days ago
view post
Post
6078
πŸ’» Data-center AI, now on a laptop: POCKET-Darwin-180B

We're releasing a 4-bit GGUF build of Darwin-180B-RSI, #1 on seven official Hugging Face leaderboards (self-reported), that runs without a GPU.

πŸ“¦ 360 GB β†’ 111 GB (4-bit GGUF, 4 files)
πŸ–₯️ No GPU: one server CPU (16 threads) at 18.4–21.0 tokens/s
πŸ’» RTX 5060 laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s
🧊 128 GB mini PC: whole model in memory, no GPU needed
🎯 MMLU-Pro, 2,000 questions, paired: original 87.65% = 4-bit 87.65%

How?
Β· Only ~3B of 180B parameters are active per token (10 of 512 experts)
Β· llama.cpp streams just the needed experts from SSD, so 32 GB RAM is enough
Β· Graft quantization: we took the proven Unsloth UD-Q4_K_XL base build and swapped in only the 300 tensors our RSI training changed (300/300 verified)

Under the hood is Model-level Recursive Self-Improvement. The model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces.

Built for teams that can't send data to an external cloud (defense, finance, public sector) to run a top-tier model fully offline.

πŸ“ Article: https://huggingface.co/blog/FINAL-Bench/data-center-ai-now-on-a-laptop-pocket-darwin-180b
πŸ€— Model: FINAL-Bench/POCKET-Darwin-180B-GGUF
🧬 Original: FINAL-Bench/Darwin-180B-RSI

#Darwin #RSI #GGUF #llamacpp #OnDevice #MoE
  • 30 replies
Β·
BoldingBuildsΒ 
posted an update about 18 hours ago
view post
Post
2279
Abliterated thinking models have a habit: they decide what to say, then keep arguing with themselves until the token budget runs out. Qwen3.8-27B Abliterated ThinkFix is an uncensored Qwen3.8-27B with one output row edited so the model closes its thinking when it's ready to answer.

On 40 sensitive prompts the edit never saw, 8k output cap: 33 replies finished cleanly instead of 22, 5 hit the limit instead of 17, 61 minutes for the set instead of 75.

Agent use was checked four ways and the edit changes nothing there. MTP draft head kept. Q4_K_M to Q8_0, each tested after the edit. The card lists what the edit was fit on and what it does not do.

BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF
  • 3 replies
Β·
DedeProGamesΒ 
posted an update 3 days ago
view post
Post
5701
how is this possible
  • 13 replies
Β·
pollixΒ 
posted an update 2 days ago
view post
Post
6172
stuntd 0.1.3 is out, and most of it started with a comment under my last post :)

@dipankarsarkar ran the support demo himself and showed my "1,000 tickets the heads never saw" were new wording of known tickets, not new kinds. On templates left out of training the heads were sure on 40% of tickets and right on only 61% of those. They also answered "what's the weather in Paris" like it was a billing ticket.

So now every head keeps the encoder vectors of what it trained on, and anything far from all of them goes to the big model no matter how confident the head is.

Screenshot is the same three requests with the gate off and on. It's local Jev mode with no provider, so the gate hands them to zero-shot Laya. Behind OpenAI or Anthropic they'd go to your model.

On the support demo it costs about 6 points of local answers on normal tickets (76.6% to 70.6%, still 97% right), and tickets from unseen templates go from 40% answered locally to 0. Funny detail: a ticket I typed by hand counted as new too, these demo heads only ever saw 20 templates.

Also in this one: auto retrain counts distinct texts and waits between runs, report shows intervals, import skips duplicates, site rm/rename, /healthz and stuntd stop.

pip install -U stuntd


https://github.com/bladedevoff/stuntd/releases/tag/v0.1.3
  • 1 reply
Β·
jialinyyzzΒ 
posted an update about 15 hours ago
view post
Post
1228
For a single-purpose 12B rewriter, how much loss per bit? Chinese breaks first.

We used llama.cpp's --kl-divergence for jialinyyzz/humanizer against bf16 weights, on our held out eval (drafts+rewrites) for English and Chinese. Standard llama-quantize (but our imatrix) without additional training. "differ" means top-choice token is not the same as what bf16 selects.

- Q6_K, 10.0 GB: KL .0030 EN / .0027 ZH. About 2 in 100 tokens differ in both.
- Q4_K_M, 7.6 GB: .0198 / .0204. About 6 in 100.
- IQ3_XXS, 4.7 GB: .138 / .175. 14 vs. 17 in 100.
- IQ2_XS, 3.8 GB: .487 / .762. 26 vs. 35 in 100.

Both languages track when you go as low as 4 bits. Below that Chinese drops off faster (Chinese KL is 1.3x English at 3 bits and 1.6x at 2 bits). Not clear why. One clue: when we added more Chinese to an imatrix (1/3 of the 800k tokens), it reduced Chinese KL by 4.7% at 2 bits. English didn't change.

(The ruler: These were running at about 8k tokens for each language. A 30690 token run agrees with English, but suggests Chinese was ~10% undercounted for 4 and 3 bits. So a bit more difference)

WIP/not released. We're distilling only the fp parts of the GGUF (block scales and norms) vs. bf16. Freezing integer codes. At 2 bits we're seeing approx. 1/2 KL (so far .487 -> .263 EN, .762 -> .319 ZH; different tokens 26 -> 20, 35 -> 22 / 100). Barely changed at 4 bits so we stopped there. No new stuff to grab yet.
Banaxi-TechΒ 
posted an update 4 days ago
view post
Post
5127
ACR 1.0 launch is being prepared and researched now!
Also I'm going to vacation tomorrow but it should still be released!

saicr
  • 3 replies
Β·
philbert440Β 
posted an update 10 days ago
view post
Post
226
Welp, I ended up getting a supermicro DGX-1 variant, so soon I'll have 8 nvlink'd V100s and 256gb of VRAM to play with, more things will be otw soon.
  • 1 reply
Β·
samuel-vitorinoΒ 
posted an update Aug 27
view post
Post
254
Sopro V2 is out: open-source voice-cloning TTS at 120M params, Apache-2.0.

- Streams with ~300 ms time-to-first-audio on a laptop CPU (0.21 RTF, and 0.07 RTF on an H100)
- English, French, German, and native European Portuguese, to my knowledge a first for open TTS
- 1.51-1.65 WER on Seed-TTS-eval test-en, competitive with models 3-14x larger (F5-TTS 1.83, CosyVoice 3 2.02, Spark-TTS 1.98)
- Zero-shot cloning from 5-20 s of reference audio
- Also runs fully in the browser (WebGPU on desktop, quantized WASM on mobile)

Try it with one command:

uvx --from sopro soprotts serve

Weights: samuel-vitorino/sopro-v2-turbo
Evals, audio samples, and how it was built: https://research.haloneuro.ai/posts/sopro-v2
Code: https://github.com/samuel-vitorino/sopro
multimodalartΒ 
posted an update Oct 16, 2025
view post
Post
39526
Want to iterate on a Hugging Face Space with an LLM?

Now you can easily convert any HF entire repo (Model, Dataset or Space) to a text file and feed it to a language model!

multimodalart/repo2txt
  • 3 replies
Β·