Skip to content
Topic

#Benchmarks

19 articles on Benchmarks — news, releases, guides and analysis from the SourceFeed engine.

One Vendor Finally Published the Attacks It Can't Catch
Article 1w ago 0

One Vendor Finally Published the Attacks It Can't Catch

An open benchmark with a public failure list beats unverifiable detection claims — but a self-scored 99.8% still isn't proof.

Ji-ho Choi
The Red Queen Comes for Your Static Evals

The Red Queen Comes for Your Static Evals

Article · 1w ago0
Models Aren't Getting Dumber — They're Getting Unbundled

Models Aren't Getting Dumber — They're Getting Unbundled

Article · 1w ago0
The New Frontier Models Are Amnesiacs by Design

The New Frontier Models Are Amnesiacs by Design

Article · 1w ago5
Z.ai Caught Up on Finding Bugs, Not on Exploiting Them

Z.ai Caught Up on Finding Bugs, Not on Exploiting Them

Article · 2w ago3
Bullet Is Fast, but It's Built on Rented Land

Bullet Is Fast, but It's Built on Rented Land

Article · 2w ago2
Anthropic's New Benchmark Scores Reasoning Nobody Can Verify

Anthropic's New Benchmark Scores Reasoning Nobody Can Verify

Article · 2w ago5
There Is No Best Language for Coding Agents

There Is No Best Language for Coding Agents

Article · 2w ago2
Intel's 18A Just Killed the ARM Efficiency Myth

Intel's 18A Just Killed the ARM Efficiency Myth

Article · 3w ago0
Claude Opus 5 Is Anthropic Undercutting Itself, on Purpose

Claude Opus 5 Is Anthropic Undercutting Itself, on Purpose

Article · 1mo ago2
GPT-5.6 Sol Rewrites the Economics of Agentic Coding

GPT-5.6 Sol Rewrites the Economics of Agentic Coding

Article · 1mo ago0
Beyond Bug Fixing: The Rise of Senior-Level AI Coding Benchmarks

Beyond Bug Fixing: The Rise of Senior-Level AI Coding Benchmarks

Article · 1mo ago2
GLM 5.2 Beats Claude on Cyber Benchmarks

GLM 5.2 Beats Claude on Cyber Benchmarks

Article · 1mo ago2
The Open-Weights Gap Depends on What You Measure

The Open-Weights Gap Depends on What You Measure

Article · 2mos ago5
Why GLM-5.2’s Low Hallucination Rate Upends the Enterprise LLM Stack

Why GLM-5.2’s Low Hallucination Rate Upends the Enterprise LLM Stack

Article · 2mos ago2
GLM-5.2 Claims Top Open-Weights Spot on Artificial Analysis

GLM-5.2 Claims Top Open-Weights Spot on Artificial Analysis

Article · 2mos ago2