Friday, September 11, 2026
Coverage: 73 videos reviewed (0 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.
New today
GPT-6 Astra and Claude Fable 5.1 launched days apart; sources cite differing benchmark leads
Claude Fable 5.1 launched Sept. 1 and OpenAI's GPT-6 Astra on Sept. 3, Julian Goldie said, both with roughly 1M-token context, up to 128,000 output tokens and the same headline API price. Goldie relayed OpenAI-published results favoring Astra: Frontier Math tier four 97.6% versus Fable's 87.8, computer use 92.7 versus 87.3, automation 41.4 versus 31.4. He also relayed Artificial Analysis figures with Fable 5.1 ahead: index 66 versus 61, and 65% versus 57.2% on Humanity's Last Exam with tools. Letta's speaker called the two very similar on the index, and Theo, reading launch notes, cited Terminal Bench Science: Astra on low 54.3 at $11, Fable 5.1 on XH high 50%. An IBM panelist cited a 95.9% Astra score on a CAD-code benchmark, versus GPT 5.6 in the 80s.
- Evidence: 0 first-party, 0 hands-on, 4 relaying
- Disagreements: Goldie relays Artificial Analysis index of 66 for Fable 5.1 versus 61 for Astra, while Letta's speaker described the two as very similar on the same index. OpenAI-published benchmarks favor Astra, while Artificial Analysis figures favor Fable 5.1; they measure different benchmarks.
- Watch: Theo - t3.gg: Fable Vs Astra Debate Is Over; Julian Goldie: GPT-6 Astra vs Claude Fable 5.1: Who Wins? (high hype)
OpenAI reportedly claims Navier-Stokes result from 10,000 agents; cost figures differ
Panelists and creators in Sept. 11 videos said OpenAI claims to have solved the Navier-Stokes Millennium Prize problem using about 10,000 agents in parallel. David Shapiro's panelist said the run began Sept. 1, lasted 88 hours, used 130 billion tokens and cost about $6.5 million, using an unannounced model stronger than Astra. Fireship cited $20 million of compute and IBM's panel about $15 million, with 17 hours of Lean verification. Matt Wolfe relayed that OpenAI said an internal model significantly more capable than Astra was used. Panelists said the proposal still requires validation by mathematicians; none of the speakers verified the claim.
- Evidence: 0 first-party, 0 hands-on, 4 relaying
- Disagreements: Compute cost is given as about $6.5 million (Shapiro's panelist), about $15 million (IBM panel) and $20 million (Fireship); IBM's panel attributes the run to Astra while Shapiro's panelist and Wolfe say an unreleased stronger model was used. Correctness is not community-verified.
- Watch: IBM Technology: OpenAI talks GPT-6 Astra and Millenium Prize, researchers create WeWor; David Shapiro: Opening act of the Singularity
Theo finds Astra ahead on 3D rendering and speed, Fable 5.1 on mergeable PRs
Theo, in a sponsored video, reported hands-on comparisons of GPT-6 Astra and Claude Fable 5.1. In his 3D game demos Astra's output looked much better, while Fable had better animation, camera and control feel. He said Astra with Codex computer use is much faster, partly because of Codex improvements on macOS. On his own pull requests Fable 5.1 needed an average of two follow-ups before merge and Astra about six, on what he called a vibe-based chart. A Rust port of TypeScript run with 40 sub-agents rose from about 30% to over 80% of the TypeScript test suite in about three days with Astra, versus about 30% with 5.6 Soul, then stalled at 82.6%. He also said Astra ignored an instruction to reuse UI code in a ping.gg rewrite.
- Evidence: 0 first-party, 1 hands-on, 0 relaying
- Watch: Theo - t3.gg: Fable Vs Astra Debate Is Over
Meta launched Muse, a personal agent on web and WhatsApp in the US
Meta launched Muse, a personal agent that acts on a per-user cloud virtual machine with a browser and storage, per Julian Goldie and Matt Wolfe, available in the United States on web and WhatsApp and through iOS and Android apps, with Meta glasses later. Goldie said most people can use it free and that Meta plans a confidential VM later this year. Both said a Meta Sentinel layer approves or blocks online actions, with confirmation required before email or payments. Wolfe said it had reached number two among US apps. In his early-access test, Wolfe connected Facebook, Instagram, Gmail and calendar and had Muse audit his AI subscriptions; it found many but missed some, including OpenAI. He called onboarding the simplest of agents he tried, with fewer integrations. David Shapiro's panelist said Zuckerberg announced Muse a couple of days earlier.
- Evidence: 0 first-party, 1 hands-on, 2 relaying
- Watch: Matt Wolfe: AI News: The AI World is REALLY Scared Right Now; Julian Goldie: NEW Meta Muse AI Agent is ABSURD! 🤯 (high hype)
DeepSeek released V4.1 Flash, an open-weights multimodal model with vendor-reported benchmarks
DeepSeek released V4.1 Flash, which Matthew Berman, reading the vendor blog, described as an open-weights 552B mixture-of-experts model (8B active for input and 16B for output, as spoken) with Terminal Bench 3.0 score 30, DeepSWE 74.2, CyberGym 88.1 and ExploitGym 15. Matt Wolfe read Artificial Analysis: V4.1 Flash 40 versus previous 36, at 27 cents per task, versus $8.75 for Fable 5 and $3.26 for GPT-6; DeepSWE 1.1 74.2 versus about 74% for Astra, Gemini 3.8 Flash and Opus 5. Sentdex said it has vision. Goldie's chart placed it near GPT 5.6 (94.1) and Claude Opus 5 (93.4) on GPQA Diamond without reading Flash's own score. DeepSeek's claimed memory savings are covered separately. Figures are vendor or relayed.
- Evidence: 0 first-party, 1 hands-on, 3 relaying
- Watch: Julian Goldie: Deepseek v4.1 is SCARY GOOD! (high hype); Sentdex: Effective Doomerism
Hands-on tests of DeepSeek V4.1 Flash show fast output but failures on harder tasks
Matthew Berman eyeballed about 200 tokens per second (a 1000-word essay in about 6 seconds), but V4.1 Flash failed his Rubik's Cube simulation in DeepSeek chat and in the Codex harness, and Paintbench and bullet-through-water tests gave mixed results. Matt Wolfe ran his SVG bench at 59 seconds and a little under two cents, and judged it weaker than GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1, which he said conflicts with its DeepSWE score. Julian Goldie built five projects in about 10 minutes in a DeepSeek harness, called quality decent but below Astra, and suggested it as a secondary model under Astra; a check in Hermes returned in about 2 seconds. Sentdex said he prefers GLM 5.x over DeepSeek V4 Flash for real work, and Berman argued cheaper open models suffice for about 95% of uses.
- Evidence: 0 first-party, 3 hands-on, 1 relaying
- Watch: Matthew Berman: Deepseek did it again...; Matt Wolfe: AI News: The AI World is REALLY Scared Right Now
Continuing stories
Also notable
- Artificial Analysis cost-per-task figures put GPT-6 Astra below Fable 5.1 and Opus 5 - Theo, relaying Artificial Analysis, said cost per task was $3.26 for Astra, almost $6 for Opus 5 and $7.60 for Fable 5.1, with Astra using about 27K tokens where Fable 5.1 used almost 80K. [0 first-party, 0 hands-on, 1 relaying] Watch: Theo - t3.gg: Fable Vs Astra Debate Is Over
- Users report GPT-6 Astra building tools and driving desktop apps in clips and demos - OpenAI-published clips in which users describe Astra: one said it built an adjustable stripe-font tool in about 15 to 20 minutes, another said it makes thumbnail adjustments in Affinity and preps and color-grades in Final Cut through computer use, and a third said it iterated website designs with matching details unprompted. [1 first-party, 0 hands-on, 1 relaying] Watch: Letta: Letta Office Hours: OpenAI's Astra Arrives on Letta
- Advantage host reports one Astra run from two photos to an uploaded print poster - In a sponsored video, The AI Advantage host said he gave GPT-6 Astra (ChatGPT desktop Work tab, medium setting) one brief, and it made and edited an image, laid out an editable Canva poster, exported a print PDF, uploaded an 18x24 in poster to a printer, then stopped the OBS recording. [0 first-party, 1 hands-on, 0 relaying] Watch: The AI Advantage: ChatGPT Images 2.5 Is Here. Together With Astra It’s Crazy
- Buckmaster and Alpige report Euler blow-up; dispute with OpenAI over credit; Tao comments - Fireship said NYU professor Tristan Buckmaster and Levent Alpige, who works at Anthropic, used Claude Code and Codex from mid-August and on Aug. [0 first-party, 0 hands-on, 1 relaying] Watch: Fireship: OpenAI's biggest math breakthrough is getting ugly...
- Jacob Cox posted Anthropic resignation warning of AI extinction risk; critics dispute framing - A pretraining researcher identified as Jacob Cox posted on Sept. [0 first-party, 0 hands-on, 3 relaying] Watch: Sentdex: Effective Doomerism
- Sentdex says Irregular ran the Hugging Face agent-hacking benchmark in a weak sandbox - Sentdex said later information showed the benchmark run in which agents hacked Hugging Face was run by a third party called Irregular, not OpenAI, and that the sandbox had internet access and apparently was just a Docker container. [0 first-party, 0 hands-on, 1 relaying] Watch: Sentdex: Effective Doomerism
- DeepSeek says V4 Pro requests redirect to V4.1 Flash from Sept. 14 - Julian Goldie relayed that DeepSeek will retire V4 Pro, with requests redirected to V4.1 Flash on Sept. [0 first-party, 0 hands-on, 1 relaying] Watch: Matthew Berman: Deepseek did it again...
- DeepSeek says V4.1 needs a quarter of KV-cache HBM and an eighth of SSD - Matthew Berman relayed DeepSeek's claim that V4.1's KV cache needs one fourth of the HBM and one eighth of the SSD storage, with memory footprint described as 8x smaller than V3.2, 13x than V4 Flash and another 4x from V4 to V4.1, as spoken. [0 first-party, 0 hands-on, 1 relaying] Watch: Matthew Berman: Deepseek did it again...
- DeepSeek V4.1 Flash API has peak and off-peak prices; free Token Harbor access reported - Matthew Berman cited DeepSeek V4.1 Flash API prices of 15 cents per million uncached input tokens off-peak and 30 cents at peak, with cached input a fraction of a penny and output at 60 cents off-peak, with a peak figure of $120 per million as spoken. [0 first-party, 1 hands-on, 1 relaying] Watch: Julian Goldie: How to use DeepSeek V4.1 Flash API for FREE!
- Cognition released SWE-2, post-trained from Kimi K3, citing its own benchmark gains - Cognition released SWE-2, post-trained from Kimi K3 with reinforcement learning, according to AI Code King, which relayed Cognition's figures: Frontier Code 11 main 44.2% (K3) to 50% (SWE-2), Terminal Bench 2.1 88.3% to 92.8%, and Terminal Bench 4 27.3% versus 57.9% for GPT-6 Astra. [0 first-party, 0 hands-on, 1 relaying] Watch: AI Code King: SWE-2 (Fully Tested): WHAT? IT ACTUALLY BEATS ASTRA & FABLE!
Models & learning
- AI Code King scored SWE-2 67/80 on KingBench 3; it asks many clarifying questions - On his KingBench 3, AI Code King scored SWE-2 at 67 of 80 (83.75%) versus 65 of 80 for DeepSeek V4.1 Flash, across eight tasks scored out of 10: SWE-2 won three, lost two and tied three. [0 first-party, 1 hands-on, 0 relaying] Watch: AI Code King: SWE-2 (Fully Tested): WHAT? IT ACTUALLY BEATS ASTRA & FABLE!
- Edge0 preview claims Qwen 3.5 35B A3B runs in under 3 GB on Macs - Fahd Mirza relayed repo claims that Edge0, a preview for macOS Apple silicon only, runs a 35B mixture-of-experts tier (Qwen 3.5, 256 experts, 4 active per token) in about 2.9 GB peak at 15 to 18 tokens per second on a Mac Mini M4 Pro, and a smaller 8B tier at 24 to 25 tokens per second in about 1 GB. [0 first-party, 0 hands-on, 1 relaying] Watch: Fahd Mirza: Run 35B Model on Phone Under 3GB Memory with Edge0
- OUI-1 diffusion UI generator tested locally: under-second screens but parser errors on dense layouts - Fahd Mirza described OUI-1 as a diffusion model fine-tuned from Google DiffusionGemma with 4B active parameters that outputs OpenUI Lang and generates a whole screen in one shot. [0 first-party, 1 hands-on, 0 relaying] Watch: Fahd Mirza: OUI-1: Builds UI Screens Instantly Locally
- LeVJEPA trains a video JEPA with one encoder and no EMA teacher, presenter reports - A presenter in a Cohere series described LeVJEPA, which uses one encoder with an MSE plus SIGReg loss and no EMA teacher or stop-gradient. [0 first-party, 1 hands-on, 0 relaying] Watch: Cohere: Lukas Kuhn - LeVJEPA Efficient & Scalable Video Pretraining without t
- Open-weight voice agent identified language correctly in 446 of about 500 traces - A speaker in a Cohere session reviewed roughly 500 traces from about 500 users over about a month, saying language was identified correctly for 446 (about 90%), end-to-end turn completion was about 89%, and 410 traces had correctly observed answer audio. [0 first-party, 1 hands-on, 0 relaying] Watch: Cohere: Suneel Sunkara - Building Voice Agents for Asian Languages Applied L
- Hugging Face shows OpenEnv CLI and GRPO training demos on small models - Hugging Face speakers described an OpenEnv CLI with init, push, pull and fork commands, validate and discover due in the next release, and about 4,000 environments on the Hub; environments are Docker apps deployable to Spaces, sandboxes, Modal, Daytona, Kubernetes or local. [1 first-party, 1 hands-on, 0 relaying] Watch: Hugging Face: Training Agents 4: From reward functions to environments.
- Hands-on tests find ChatGPT Images 2.5 edits more consistent but not perfect - The AI Advantage host edited a mug to forest green and a poster from coffee to croissant; framing stayed nearly identical but lighting and table structure shifted, and a multi-edit chain on a headshot kept identity. [0 first-party, 1 hands-on, 1 relaying] Watch: The AI Advantage: ChatGPT Images 2.5 Is Here. Together With Astra It’s Crazy
- Hermes Desktop manages llama.cpp and picks local model builds per machine - Julian Goldie said Hermes Desktop downloads and manages llama.cpp, chooses a model build to fit the machine, handles context and GPU layers automatically, and needs no account or API key; models can also come from Hugging Face or a local GGUF. [0 first-party, 0 hands-on, 1 relaying] Watch: Julian Goldie: Hermes Desktop Can Now Set Up Local AI in ONE Click