@DedeProGames your response?
Banaxi PRO
Banaxi-Tech
AI & ML interests
SLMs, training from scratch, LoRA, TTS, Ternary models. AI Interpretability. BCI. Contact at banaxitech@gmail.com
Recent Activity
new activity about 11 hours ago
bench-labs/cagliostro-v3.5:bruh commentedon their article about 11 hours ago
Some OrionLLM Models are NOT legitimate liked a model about 11 hours ago
bench-labs/cagliostro-v3.5Organizations
bruh
11
#1 opened about 12 hours ago
by
Banaxi-Tech
commented on Some OrionLLM Models are NOT legitimate about 11 hours ago
Add cagliostro-v3.5
4
#183 opened about 12 hours ago
by
TobiasLogic
upvoted a collection about 18 hours ago
Banaxi-Tech, look at the new UI
3
#5 opened 1 day ago
by
DedeProGames
replied to their post 2 days ago
Italy!
commented on Some OrionLLM Models are NOT legitimate 2 days ago
The hashes also the same here
commented on Some OrionLLM Models are NOT legitimate 2 days ago
@BananaMindBot is this legal?
upvoted an article 2 days ago
Article
Some OrionLLM Models are NOT legitimate
Banaxi-Tech
โข โข 5published an article 2 days ago
Article
Some OrionLLM Models are NOT legitimate
Banaxi-Tech
โข โข 5replied to SeaWolf-AI's post 3 days ago
they just rremoved them this post was full of em dashes
reacted to SeaWolf-AI's post with ๐ 3 days ago
Post
6023
๐ป Data-center AI, now on a laptop: POCKET-Darwin-180B
We're releasing a 4-bit GGUF build of Darwin-180B-RSI, #1 on seven official Hugging Face leaderboards (self-reported), that runs without a GPU.
๐ฆ 360 GB โ 111 GB (4-bit GGUF, 4 files)
๐ฅ๏ธ No GPU: one server CPU (16 threads) at 18.4โ21.0 tokens/s
๐ป RTX 5060 laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s
๐ง 128 GB mini PC: whole model in memory, no GPU needed
๐ฏ MMLU-Pro, 2,000 questions, paired: original 87.65% = 4-bit 87.65%
How?
ยท Only ~3B of 180B parameters are active per token (10 of 512 experts)
ยท llama.cpp streams just the needed experts from SSD, so 32 GB RAM is enough
ยท Graft quantization: we took the proven Unsloth UD-Q4_K_XL base build and swapped in only the 300 tensors our RSI training changed (300/300 verified)
Under the hood is Model-level Recursive Self-Improvement. The model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces.
Built for teams that can't send data to an external cloud (defense, finance, public sector) to run a top-tier model fully offline.
๐ Article: https://huggingface.co/blog/FINAL-Bench/data-center-ai-now-on-a-laptop-pocket-darwin-180b
๐ค Model: FINAL-Bench/POCKET-Darwin-180B-GGUF
๐งฌ Original: FINAL-Bench/Darwin-180B-RSI
#Darwin #RSI #GGUF #llamacpp #OnDevice #MoE
We're releasing a 4-bit GGUF build of Darwin-180B-RSI, #1 on seven official Hugging Face leaderboards (self-reported), that runs without a GPU.
๐ฆ 360 GB โ 111 GB (4-bit GGUF, 4 files)
๐ฅ๏ธ No GPU: one server CPU (16 threads) at 18.4โ21.0 tokens/s
๐ป RTX 5060 laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s
๐ง 128 GB mini PC: whole model in memory, no GPU needed
๐ฏ MMLU-Pro, 2,000 questions, paired: original 87.65% = 4-bit 87.65%
How?
ยท Only ~3B of 180B parameters are active per token (10 of 512 experts)
ยท llama.cpp streams just the needed experts from SSD, so 32 GB RAM is enough
ยท Graft quantization: we took the proven Unsloth UD-Q4_K_XL base build and swapped in only the 300 tensors our RSI training changed (300/300 verified)
Under the hood is Model-level Recursive Self-Improvement. The model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces.
Built for teams that can't send data to an external cloud (defense, finance, public sector) to run a top-tier model fully offline.
๐ Article: https://huggingface.co/blog/FINAL-Bench/data-center-ai-now-on-a-laptop-pocket-darwin-180b
๐ค Model: FINAL-Bench/POCKET-Darwin-180B-GGUF
๐งฌ Original: FINAL-Bench/Darwin-180B-RSI
#Darwin #RSI #GGUF #llamacpp #OnDevice #MoE