AI news, roughly super. A daily briefing from what the AI YouTube world actually said.

Friday, September 11, 2026

Coverage: 73 videos reviewed (0 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.

New today

GPT-6 Astra and Claude Fable 5.1 launched days apart; sources cite differing benchmark leads
Claude Fable 5.1 launched Sept. 1 and OpenAI's GPT-6 Astra on Sept. 3, Julian Goldie said, both with roughly 1M-token context, up to 128,000 output tokens and the same headline API price. Goldie relayed OpenAI-published results favoring Astra: Frontier Math tier four 97.6% versus Fable's 87.8, computer use 92.7 versus 87.3, automation 41.4 versus 31.4. He also relayed Artificial Analysis figures with Fable 5.1 ahead: index 66 versus 61, and 65% versus 57.2% on Humanity's Last Exam with tools. Letta's speaker called the two very similar on the index, and Theo, reading launch notes, cited Terminal Bench Science: Astra on low 54.3 at $11, Fable 5.1 on XH high 50%. An IBM panelist cited a 95.9% Astra score on a CAD-code benchmark, versus GPT 5.6 in the 80s.

OpenAI reportedly claims Navier-Stokes result from 10,000 agents; cost figures differ
Panelists and creators in Sept. 11 videos said OpenAI claims to have solved the Navier-Stokes Millennium Prize problem using about 10,000 agents in parallel. David Shapiro's panelist said the run began Sept. 1, lasted 88 hours, used 130 billion tokens and cost about $6.5 million, using an unannounced model stronger than Astra. Fireship cited $20 million of compute and IBM's panel about $15 million, with 17 hours of Lean verification. Matt Wolfe relayed that OpenAI said an internal model significantly more capable than Astra was used. Panelists said the proposal still requires validation by mathematicians; none of the speakers verified the claim.

Theo finds Astra ahead on 3D rendering and speed, Fable 5.1 on mergeable PRs
Theo, in a sponsored video, reported hands-on comparisons of GPT-6 Astra and Claude Fable 5.1. In his 3D game demos Astra's output looked much better, while Fable had better animation, camera and control feel. He said Astra with Codex computer use is much faster, partly because of Codex improvements on macOS. On his own pull requests Fable 5.1 needed an average of two follow-ups before merge and Astra about six, on what he called a vibe-based chart. A Rust port of TypeScript run with 40 sub-agents rose from about 30% to over 80% of the TypeScript test suite in about three days with Astra, versus about 30% with 5.6 Soul, then stalled at 82.6%. He also said Astra ignored an instruction to reuse UI code in a ping.gg rewrite.

Meta launched Muse, a personal agent on web and WhatsApp in the US
Meta launched Muse, a personal agent that acts on a per-user cloud virtual machine with a browser and storage, per Julian Goldie and Matt Wolfe, available in the United States on web and WhatsApp and through iOS and Android apps, with Meta glasses later. Goldie said most people can use it free and that Meta plans a confidential VM later this year. Both said a Meta Sentinel layer approves or blocks online actions, with confirmation required before email or payments. Wolfe said it had reached number two among US apps. In his early-access test, Wolfe connected Facebook, Instagram, Gmail and calendar and had Muse audit his AI subscriptions; it found many but missed some, including OpenAI. He called onboarding the simplest of agents he tried, with fewer integrations. David Shapiro's panelist said Zuckerberg announced Muse a couple of days earlier.

DeepSeek released V4.1 Flash, an open-weights multimodal model with vendor-reported benchmarks
DeepSeek released V4.1 Flash, which Matthew Berman, reading the vendor blog, described as an open-weights 552B mixture-of-experts model (8B active for input and 16B for output, as spoken) with Terminal Bench 3.0 score 30, DeepSWE 74.2, CyberGym 88.1 and ExploitGym 15. Matt Wolfe read Artificial Analysis: V4.1 Flash 40 versus previous 36, at 27 cents per task, versus $8.75 for Fable 5 and $3.26 for GPT-6; DeepSWE 1.1 74.2 versus about 74% for Astra, Gemini 3.8 Flash and Opus 5. Sentdex said it has vision. Goldie's chart placed it near GPT 5.6 (94.1) and Claude Opus 5 (93.4) on GPQA Diamond without reading Flash's own score. DeepSeek's claimed memory savings are covered separately. Figures are vendor or relayed.

Hands-on tests of DeepSeek V4.1 Flash show fast output but failures on harder tasks
Matthew Berman eyeballed about 200 tokens per second (a 1000-word essay in about 6 seconds), but V4.1 Flash failed his Rubik's Cube simulation in DeepSeek chat and in the Codex harness, and Paintbench and bullet-through-water tests gave mixed results. Matt Wolfe ran his SVG bench at 59 seconds and a little under two cents, and judged it weaker than GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1, which he said conflicts with its DeepSWE score. Julian Goldie built five projects in about 10 minutes in a DeepSeek harness, called quality decent but below Astra, and suggested it as a secondary model under Astra; a check in Hermes returned in about 2 seconds. Sentdex said he prefers GLM 5.x over DeepSeek V4 Flash for real work, and Berman argued cheaper open models suffice for about 95% of uses.

Continuing stories

Also notable

Models & learning