Read article
Wellfunded

Benchmarks

a16z Backs Vals, AI's Independent Model Scorekeeper

Aug 14, 2026

a16z Backs Vals, AI's Independent Model Scorekeeper

Andreessen Horowitz has invested in Vals, a startup building real-world AI benchmarks for high-stakes professional domains like law, finance, and coding. Unlike academic leaderboards that get gamed or saturated, Vals tests models on actual workflows — and deliberately retires benchmarks once frontier models max them out. a16z sees the company as AI's emerging equivalent of Moody's or UL certification: the independent trust layer enterprise buyers need as agentic AI moves into consequential work.

AI

Venture Capital

Startups

Hark Unveils Handoff, a Cheap Computer-Use Agent

Aug 6, 2026

Hark Unveils Handoff, a Cheap Computer-Use Agent

Brett Adcock's secretive AI startup Hark has launched Handoff, a computer-use agent claiming top scores on web navigation benchmarks at a fraction of frontier model pricing. But the comparisons leave out the newest models from OpenAI and Anthropic — raising questions the company hasn't yet answered.

AI

Startups

Computer Use Agents

Qwen3.8-Max Tops Agentic Benchmarks, May Go Open Weight

Aug 4, 2026

Qwen3.8-Max Tops Agentic Benchmarks, May Go Open Weight

Alibaba's Qwen team has released Qwen3.8-Max, a 2.4-trillion-parameter MoE model that outscores GPT-5.6 Sol Max and Fable 5 on OSWorld-Verified. With open weights potentially arriving next week and pricing well below leading U.S. rivals, the model could reshape enterprise AI adoption.

AI

Enterprise AI

Open Source AI

ChatGPT Nears 1 Billion Weekly Users, Behind Schedule

Jul 31, 2026

ChatGPT Nears 1 Billion Weekly Users, Behind Schedule

ChatGPT is approaching 1 billion weekly active users, a milestone OpenAI originally targeted for late 2024. The figure arrives amid a busy week for the company: new transcription models, a security-focused Codex CLI, and a benchmark controversy around its o3 successor, Sol.

AI

OpenAI

Startups

Hugging Face's VoiceEQ Benchmarks Human-Like Voice AI

Jul 15, 2026

Hugging Face's VoiceEQ Benchmarks Human-Like Voice AI

Hugging Face has launched Real World VoiceEQ, a benchmark designed to evaluate the human quality of voice AI systems beyond basic transcription accuracy. The leaderboard includes audio samples and aims to give developers a more meaningful way to compare speech models in real-world conditions.

Voice AI

Benchmarks

AI