Skip to content
Topic

#Quantization

13 articles on Quantization — news, releases, guides and analysis from the SourceFeed engine.

llmfit does the local-LLM math you've been faking
Article 1w ago 6

llmfit does the local-LLM math you've been faking

The k8sgpt creator's Rust CLI right-sizes models to your GPU — keeping its catalog honest is the hard part.

Mariana Souza
Convert and Quantize Hugging Face Models to GGUF for llama.cpp

Convert and Quantize Hugging Face Models to GGUF for llama.cpp

Tutorial · 2w ago0
Latent Reasoning Escapes the Lab, Bolted Onto DeepSeek-V4

Latent Reasoning Escapes the Lab, Bolted Onto DeepSeek-V4

Article · 2w ago2
Run Mixtral 8x7B Locally with llama.cpp and Benchmark MoE vs. Dense

Run Mixtral 8x7B Locally with llama.cpp and Benchmark MoE vs. Dense

Tutorial · 3w ago0
An 8B Fine-Tune Now Fits in 4 GB of VRAM

An 8B Fine-Tune Now Fits in 4 GB of VRAM

Article · 3w ago1
Your SSD Is the New VRAM for Local LLMs

Your SSD Is the New VRAM for Local LLMs

Article · 3w ago1
Cloudflare's Quantization Math Is Right. The Disclosure Isn't

Cloudflare's Quantization Math Is Right. The Disclosure Isn't

Article · 3w ago0
Quantize the Decode, Not the Prefill

Quantize the Decode, Not the Prefill

Article · 3w ago5
An LLM on an $8 Chip Is Mostly a Memory Trick

An LLM on an $8 Chip Is Mostly a Memory Trick

Article · 1mo ago0
Bonsai 27B Puts Real Agents on Phones

Bonsai 27B Puts Real Agents on Phones

Article · 1mo ago1
Quantize and Run Llama 3.2 on Apple Silicon with llama.cpp

Quantize and Run Llama 3.2 on Apple Silicon with llama.cpp

Tutorial · 2mos ago0
Demystifying Integer Quantization for Neural Network Inference

Demystifying Integer Quantization for Neural Network Inference

Article · 2mos ago0
Xiaomi's MiMo-V2.5-Pro-UltraSpeed Pushes a 1T Model Past 1000 Tokens/Sec on Commodity GPUs

Xiaomi's MiMo-V2.5-Pro-UltraSpeed Pushes a 1T Model Past 1000 Tokens/Sec on Commodity GPUs

News · 2mos ago5