Skip to content
Topic

#Inference

23 articles on Inference — news, releases, guides and analysis from the SourceFeed engine.

Inference ASICs Won the Benchmarks, Then Lost Their Independence
Article 1w ago 2

Inference ASICs Won the Benchmarks, Then Lost Their Independence

AMD bought Taalas, Nvidia gutted Groq, Cerebras went public — how to place inference workloads after the shakeout.

Mariana Souza
llmfit does the local-LLM math you've been faking

llmfit does the local-LLM math you've been faking

Article · 1w ago6
Mistral 3 Was Never About the Leaderboard

Mistral 3 Was Never About the Leaderboard

Article · 2w ago2
Nvidia's Switchyard Finally Makes Model Routing Boring

Nvidia's Switchyard Finally Makes Model Routing Boring

Article · 2w ago1
Nvidia's Real Product Here Is the Router, Not the Model

Nvidia's Real Product Here Is the Router, Not the Model

Article · 2w ago1
NVIDIA's Switchyard Matters More Than Its New Model

NVIDIA's Switchyard Matters More Than Its New Model

Article · 2w ago4
ECS Finally Learns to Schedule a Slice of a GPU

ECS Finally Learns to Schedule a Slice of a GPU

Article · 3w ago0
Your SSD Is the New VRAM for Local LLMs

Your SSD Is the New VRAM for Local LLMs

Article · 3w ago1
Your SSD Is the New VRAM

Your SSD Is the New VRAM

Article · 3w ago1
Cloudflare's Quantization Math Is Right. The Disclosure Isn't

Cloudflare's Quantization Math Is Right. The Disclosure Isn't

Article · 3w ago0
AirLLM's 4GB 70B Trick Is Real, and Beside the Point

AirLLM's 4GB 70B Trick Is Real, and Beside the Point

Article · 3w ago1
A 26B Model in 2 GB of RAM, Courtesy of Your SSD

A 26B Model in 2 GB of RAM, Courtesy of Your SSD

Article · 1mo ago2
Google Split Its TPU in Two Because Agents Broke Inference

Google Split Its TPU in Two Because Agents Broke Inference

Article · 1mo ago2
Discarded Teslas Still Deliver Local AI VRAM

Discarded Teslas Still Deliver Local AI VRAM

Article · 1mo ago0
Popping the CPU-GPU Latency Bubble in Inference

Popping the CPU-GPU Latency Bubble in Inference

Article · 1mo ago2
OpenAI Jalapeno and the Shift to Custom Inference Silicon

OpenAI Jalapeno and the Shift to Custom Inference Silicon

Article · 2mos ago7