Read article
Wellfunded

AI Safety

Why Frontier AI Still Hallucinates, and What to Do

Jul 11, 2026

Why Frontier AI Still Hallucinates, and What to Do

Even the most advanced AI models continue to fabricate facts, citations, and events with alarming confidence. This deep-dive explores recent hallucination failures and unpacks the architectural reasons they're so hard to eliminate.

Hallucination

Large Language Models

AI Reliability

Inside the Black Box: How the US Cleared OpenAI's Model

Jul 10, 2026

Inside the Black Box: How the US Cleared OpenAI's Model

The U.S. government signed off on OpenAI's latest frontier model — but the details of that safety review remain murky. What criteria were used, who was involved, and what does this mean for AI oversight going forward?

OpenAI

AI Safety

Government Oversight

Anthropic's J-Lens Finds a Hidden 'Workspace' in Claude

Jul 7, 2026

Anthropic's J-Lens Finds a Hidden 'Workspace' in Claude

Anthropic researchers have discovered that Claude spontaneously developed an internal structure mirroring global workspace theory — a leading neuroscientific model of human consciousness. The finding, built around a new interpretability tool called the J-lens, has already begun reshaping how Anthropic monitors Claude for safety risks, including detecting silent strategic reasoning that never surfaces in the model's output.

Anthropic

Claude

AI Safety

Anthropic Strikes Deal to Restore Claude Fable 5 Access

Jul 1, 2026

Anthropic Strikes Deal to Restore Claude Fable 5 Access

Anthropic reached an agreement with the Trump administration to extend a key safety guardrail on its Claude Fable 5 model, prompting the Commerce Department to lift export controls that had effectively taken the model offline. The fix targets a specific jailbreak behavior flagged in an Amazon research paper — but the company's troubles with the Pentagon aren't fully resolved.

Anthropic

AI Safety

AI Policy

Trump Admin Lifts Export Curbs on Anthropic's AI Models

Jul 1, 2026

Trump Admin Lifts Export Curbs on Anthropic's AI Models

The Commerce Department is removing export restrictions on Anthropic's most powerful AI models after the company struck a deal to strengthen safety safeguards. The move follows weeks of behind-the-scenes negotiations — and a notable shift in how Anthropic engaged with the administration.

Anthropic

AI regulation

export controls

White House Pressures OpenAI to Delay GPT-5.6

Jun 26, 2026

White House Pressures OpenAI to Delay GPT-5.6

The Trump administration has reportedly asked OpenAI to hold back its latest model, GPT-5.6, from a broad public launch. Instead, OpenAI plans to share the model with a limited group of select partners while it navigates the White House's safety concerns.

OpenAI

GPT-5

AI Safety

Patronus AI Raises $50M for Simulated 'Digital Worlds'

Jun 26, 2026

Patronus AI Raises $50M for Simulated 'Digital Worlds'

Patronus AI, founded by former Meta AI researchers, has closed a $50M funding round to develop simulated environments that rigorously stress-test AI agents. Investor demand signals the market's growing urgency around AI reliability and safety infrastructure.

AI Agents

Funding

AI Safety