Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
nawoa lanor
nawoalanor
27
1
1
Follow
danielrmay's profile picture
webbrain-one-853887's profile picture
2 followers
·
2 following
AI & ML interests
None yet
Recent Activity
new
activity
5 days ago
unsloth/Qwen3.8-Flash-Next-GGUF:
llama.cpp now supports MTP in main, but these quants don't seem compatible
new
activity
5 days ago
unsloth/Qwen3.8-Flash-Next-GGUF:
Does this quant actually support engram SSD offload or not in llama.cpp??
new
activity
5 days ago
unsloth/Qwen3.8-Flash-Next-GGUF:
Any way to separately select ngram quantization and model quantization?
View all activity
Organizations
None yet
nawoalanor
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
unsloth/Qwen3.8-Flash-Next-GGUF
5 days ago
llama.cpp now supports MTP in main, but these quants don't seem compatible
5
#79 opened 9 days ago by
FlorinAndrei
Does this quant actually support engram SSD offload or not in llama.cpp??
3
#76 opened 14 days ago by
EnderOLED
Any way to separately select ngram quantization and model quantization?
1
#81 opened 5 days ago by
nawoalanor
New activity in
unsloth/Qwen3.8-Flash-Next-GGUF
18 days ago
Request: a UD quant between IQ4_XS (93.7 GB) and Q4_K_XL (111 GB) for 128 GB unified-memory machines (Strix Halo)
👍
2
4
#67 opened about 1 month ago by
bitlamas
Run 170 tokens/s Qwen3.8-Flash with MTP! ⚡
🔥
❤️
13
25
#56 opened about 1 month ago by
danielhanchen
New activity in
zai-org/GLM-5.3-Flash
about 2 months ago
Could've kept the active params 9b or below.
👍
1
2
#15 opened about 2 months ago by
cinnybun02
WE NEED 35B A3B
👍
🔥
19
3
#3 opened about 2 months ago by
AsThirtyThree
How is this "Flash"?
👍
🔥
25
20
#10 opened about 2 months ago by
MasterYoba
New activity in
magnitudedev/Qwen3.8-27B-DSpark-GGUF
about 2 months ago
Performance results
5
#1 opened about 2 months ago by
nawoalanor
New activity in
erlidev/Qwen3.8-27B-DSpark-GGUF
about 2 months ago
Speedup results
👍
1
#1 opened about 2 months ago by
nawoalanor
New activity in
unsloth/MiniMax-M2.7-GGUF
4 months ago
Do NOT use CUDA 13.2
🤝
👍
12
1
#2 opened 6 months ago by
danielhanchen
New activity in
deepseek-ai/DeepSeek-V4-Flash
4 months ago
Too big to run locally.
🤯
👍
12
21
#12 opened 6 months ago by
Dampfinchen
New activity in
MiniMaxAI/MiniMax-M3
4 months ago
427B params? This is not intelligence, its brute force.
😔
👀
4
17
#6 opened 4 months ago by
Nerdsking
New activity in
darkc0de/XORTRON.CriminalComputing.LARGE.2026.3
4 months ago
Abnormally poor spelling?
6
#4 opened 5 months ago by
nawoalanor
New activity in
MiniMaxAI/MiniMax-M3
4 months ago
Please release official lower-bit QAT versions
➕
❤️
12
2
#3 opened 4 months ago by
nawoalanor
New activity in
canada-quant/DeepSeek-V4-Flash-W4A16-FP8-MTP
4 months ago
vLLM compressed-tensors loader fails on shared_experts.gate_up_proj — missing quant target for fused shared-expert projection?
3
#1 opened 4 months ago by
ajdoosh
New activity in
nvidia/MiniMax-M2.7-NVFP4
5 months ago
licence?
1
#1 opened 6 months ago by
celikburak
New activity in
MiniMaxAI/MiniMax-M2.7
6 months ago
This LLM is a test maxer, not a general purpose AI model.
👍
3
21
#17 opened 6 months ago by
phil111
Need a 5% and 15% REAP
7
#10 opened 6 months ago by
nawoalanor
GLM-5.1 open ! >>> MiniMax-M2.7 open ! >>> Qwen 3.6 closed !
🚀
2
5
#14 opened 6 months ago by
Ai-Robots
Load more