Skip to content
Topic

#Vllm

9 articles on Vllm — news, releases, guides and analysis from the SourceFeed engine.

Google's 4.7x Qwen 3.5 Speedup Is a Sharding Story
Article 3d ago 3

Google's 4.7x Qwen 3.5 Speedup Is a Sharding Story

Two KV heads and 512 experts pushed Ironwood onto DP+EP, the same topology vLLM now recommends for GPUs.

Priya Nair
Latent Reasoning Escapes the Lab, Bolted Onto DeepSeek-V4

Latent Reasoning Escapes the Lab, Bolted Onto DeepSeek-V4

Article · 2w ago2
Your Agents Are Waiting on the CPU, Not the GPU

Your Agents Are Waiting on the CPU, Not the GPU

Article · 3w ago5
DeepSeek V4 Flash on One AMD GPU Took Nine Patches

DeepSeek V4 Flash on One AMD GPU Took Nine Patches

Article · 3w ago1
The Real Cost of 'Just Use vLLM,' According to Netflix

The Real Cost of 'Just Use vLLM,' According to Netflix

Article · 1mo ago2
Autoscale GPU Inference on EKS with Karpenter and Spot Instances

Autoscale GPU Inference on EKS with Karpenter and Spot Instances

Tutorial · 1mo ago0
The Real Cost of Running SOTA LLMs Locally

The Real Cost of Running SOTA LLMs Locally

Article · 1mo ago4
Ornith-1.0: Coding Models That Train Their Own Agent Scaffolds

Ornith-1.0: Coding Models That Train Their Own Agent Scaffolds

Article · 2mos ago0
Serve an Open-Source LLM at Scale with vLLM on a Rented GPU Instance

Serve an Open-Source LLM at Scale with vLLM on a Rented GPU Instance

Tutorial · 2mos ago0