Skip to content
Topic

#Llm Inference

19 articles on Llm Inference — news, releases, guides and analysis from the SourceFeed engine.

Google's 4.7x Qwen 3.5 Speedup Is a Sharding Story
Article 3d ago 3

Google's 4.7x Qwen 3.5 Speedup Is a Sharding Story

Two KV heads and 512 experts pushed Ironwood onto DP+EP, the same topology vLLM now recommends for GPUs.

Priya Nair
You're Not Buying Compute, You're Buying Utilization

You're Not Buying Compute, You're Buying Utilization

Article · 1w ago1
Speed Is Now a Paid Tier at OpenAI

Speed Is Now a Paid Tier at OpenAI

Article · 2w ago1
NVIDIA's roadmap is a memory story, not a compute story

NVIDIA's roadmap is a memory story, not a compute story

Article · 2w ago0
Latent Reasoning Escapes the Lab, Bolted Onto DeepSeek-V4

Latent Reasoning Escapes the Lab, Bolted Onto DeepSeek-V4

Article · 2w ago2
Your Agents Are Waiting on the CPU, Not the GPU

Your Agents Are Waiting on the CPU, Not the GPU

Article · 3w ago5
How a 20B Model Hits 120 tok/s on an iPhone

How a 20B Model Hits 120 tok/s on an iPhone

Article · 3w ago1
DeepSeek V4 Flash on One AMD GPU Took Nine Patches

DeepSeek V4 Flash on One AMD GPU Took Nine Patches

Article · 3w ago1
Quantize the Decode, Not the Prefill

Quantize the Decode, Not the Prefill

Article · 3w ago5
Linear Attention Just Graduated to Frontier Scale

Linear Attention Just Graduated to Frontier Scale

Article · 1mo ago2
The Real Cost of 'Just Use vLLM,' According to Netflix

The Real Cost of 'Just Use vLLM,' According to Netflix

Article · 1mo ago2
Hetzner Is Quietly Commoditizing LLM Inference

Hetzner Is Quietly Commoditizing LLM Inference

Article · 1mo ago1
KTransformers Turned CPU Offloading Into Real Infrastructure

KTransformers Turned CPU Offloading Into Real Infrastructure

Article · 1mo ago1
AI's Energy Problem Isn't Your Chatbot Query

AI's Energy Problem Isn't Your Chatbot Query

Article · 1mo ago2
VRAM Beats TOPS for 2026 Local AI GPUs

VRAM Beats TOPS for 2026 Local AI GPUs

Article · 1mo ago2
Ditching HBM: Inside the Monolithic 3D AI ASIC

Ditching HBM: Inside the Monolithic 3D AI ASIC

Article · 2mos ago2