Friday, September 4, 2026
Coverage: 81 videos reviewed (1 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.
New today
Continuing stories
- OpenAI launches GPT-6 Astra with staged access, $10/$50 pricing and computer-use focus - OpenAI announced GPT-6 Astra (Sept 3) for ChatGPT, Codex and the API, initially to a limited set of organizations and early users, with paid ChatGPT tiers (reportedly incl. [1 first-party, 0 hands-on, 8 relaying] Watch: Theo - t3.gg: It's Here.
- ARC-AGI-3: Astra 99.9% with OpenAI adapter versus 62.7% in standard harness - OpenAI's launch chart showed Astra near-saturating ARC-AGI-3 (99-99.9%), with ARC Prize saying it beat the human action-efficiency baseline on 96% of levels. [0 first-party, 0 hands-on, 5 relaying] Watch: Prompt Engineering: GPT-6 Astra: The harness matters more than you think
- Astra benchmark results: FrontierMath T4, coding, OSWorld 2.0 and Erdos problems - Reported results include FrontierMath Tier 4 about 97.6-98% (GPT-5.6 Sol 83%, Fable 5.1 87.8%), Terminal Bench 4.0 57.9% vs Fable 5.1 55.8% with near-ties on Deep SWE and Frontier Code, OSWorld 2.0 ~72.6% vs Sol 65.7% at roughly half the task time, and Epoch AI saw 2 of 68 unsolved Erdos problems solved at very high repeated-attempt cost. [0 first-party, 0 hands-on, 4 relaying] Watch: AI Code King: GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model an
- OpenAI rates Astra Critical for cyber; system card flags monitorability and bio-eval concerns - OpenAI classifies Astra as its first model at the Critical cyber threshold under its Preparedness Framework, with exploit bench 100% and gated advanced exploit generation, plus a reported $1B credit offer to cyber defenders. [0 first-party, 0 hands-on, 4 relaying] Watch: AI Explained: GPT 6 Astra, so good even OpenAI are worried
- Hands-on Astra tests: strong for computer use and long tasks, mixed against Fable 5.1 - Multiple creators tested Astra: Every's team found it a strong daily driver but slightly behind Fable 5.1 on the biggest tasks, though Astra slightly edged Fable 5.1 in a 50-comparison blind writing test and saturated one clone benchmark. [0 first-party, 4 hands-on, 1 relaying] Watch: Every: VIBE CHECK: GPT-6 ASTRA
- Anthropic releases Claude Fable 5.1 and Mythos 5.1 with cache-price cut - Anthropic released Fable 5.1 (public, classifier-wrapped) and Mythos 5.1 (vetted enterprises), described as the same underlying model with two access tiers. [0 first-party, 0 hands-on, 4 relaying] Watch: Nate B Jones: Everyone's Testing Claude Fable 5.1 On Code. It Made Me A 37-Second Fi
Also notable
- Artificial Analysis rates Astra 61, tied with GPT-5.6 Sol and below Fable 5.1 - Artificial Analysis Intelligence Index gives Astra 61, equal to GPT-5.6 Sol and five below Claude Fable 5.1 (66), with cost per task reportedly about 75% higher than Sol and regressions on GDPval and other areas. [0 first-party, 0 hands-on, 3 relaying] Watch: AI Explained: GPT 6 Astra, so good even OpenAI are worried
- Fable 5.1 hands-on: long refactor, DCF workbook, Blender film and writing test - Chris Hay ran a long open-source refactor in ~30 minutes using ~20% of the weekly limit in two hours. [0 first-party, 1 hands-on, 1 relaying] Watch: Nate B Jones: Everyone's Testing Claude Fable 5.1 On Code. It Made Me A 37-Second Fi
- Meta Muse Spark 1.3 released with low pricing; tops benchmarks but weak in a hands-on test - Meta's fourth release in five months from Meta Superintelligence Labs is priced at $1.25/$4.25 per million tokens (10c/20c if data is used for training). [0 first-party, 1 hands-on, 1 relaying] Watch: Matt Wolfe: AI News: The Most Insane Week So Far This Year!
- Alibaba Qwen 3.8 Max: 2.4T-parameter MoE, 1M context, free on chat.qwen.ai - Alibaba's Qwen3.8-Max is described as a 2.4T-parameter sparse MoE (~95B active) with native 1M context and multimodal input; Alibaba claims a 16-day continuous autonomous coding demo. [0 first-party, 0 hands-on, 2 relaying] Watch: GitHub: The Download: Attach images in GitHub CLI, Qwen3.8-Max, AI pull reques
- IFM/MBZUAI releases K2 Horizon open models (0.9B-375B, Apache 2) with data and code - Institute of Foundation Models released six K2 Horizon models with Apache 2 weights, training data, checkpoints and code, including a ~375B MoE (~23B active, 512K context). [0 first-party, 1 hands-on, 1 relaying] Watch: Fahd Mirza: K2 Horizon: 0.9B, 7B, and 32B Tested Locally, Real Results
- MiniMax M3 open model: ~400B MoE with vision and 1M context via sparse attention - MiniMax guest says M3 has roughly 400-428B total/20B active parameters, native multimodal training, 1M context via MiniMax Sparse Attention, apps reaching 300M+ people, and plans for trillion-parameter open models. [0 first-party, 0 hands-on, 1 relaying] Watch: AI Engineer: Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, M
- GLM 5.3 and GLM 5.3 Flash released; Flash tested locally on RTX Pro 6000s - Zhipu released GLM 5.3 and 5.3 Flash (reported 320B in one source, 380B in another, with vision). [0 first-party, 1 hands-on, 1 relaying] Watch: Sentdex: All Roads Lead back To GLM!
- GitHub launches Project HydraFusion research preview with routing claims - HydraFusion picks single-model, cascade or critique paths; GitHub claims a Terminal Bench 2.1 win over Opus 5 at 67% lower cost (first-party claim). [1 first-party, 0 hands-on, 0 relaying] Watch: GitHub: Introducing Project HydraFusion: multi-model orchestration in GitHub C
- Meta Muse Glimmer 30B reportedly runs agentic coding loops on one 24GB GPU - Presenter says Meta released Muse Glimmer 30B; claim is secondhand. [0 first-party, 0 hands-on, 1 relaying] Watch: GitHub: The Download: Attach images in GitHub CLI, Qwen3.8-Max, AI pull reques
- Hugging Face releases 200+ WebGPU kernels, a Kernels JS library and Fleet benchmark tool - Kernels ship as Jinja templates generating WGSL; a demo ran ~60 fps vs ~6.75 fps in plain JavaScript; Fleet crowdsources GPU benchmarks. [1 first-party, 0 hands-on, 0 relaying] Watch: Hugging Face: We shipped 207 WebGPU Kernels for Browser AI
Models & learning
- Google releases Gemini 3.8 Flash for agentic loops with 1M context - Google announced Gemini 3.8 Flash on Sept 2 (third Flash in six weeks), with 1M input/65K output tokens and low/medium/high thinking, in AI Studio, API, Antigravity and Stitch. [0 first-party, 0 hands-on, 2 relaying] Watch: Julian Goldie: New Google AI Studio Update Is WILD!
- Third-party and user reports on Astra: Pokemon, AutomationBench, legal, long agent runs - Relayed reports: Astra finishes Pokemon Fire Red in 18 hours (Sol 96), beats Fable 5.1 on Zapier AutomationBench, a legal benchmark rose from 69% to 93% on NDA review, Ethan Mollick ran it ~5 days on an email wiki, a user claims 55 agents audited 10 financial models, Dan Shipper calls it the best writing model, and sponsor Box reports a 3% gain on its eval. [0 first-party, 0 hands-on, 2 relaying] Watch: The AI Advantage: GPT-6 Astra: 20 Real Examples From Useful to Almost Impossible
- Qwen 3.8 27B with thinking on exhausted its output budget without writing a file - A tester found the same build task worked in about 5 minutes with thinking off but failed with thinking on. [0 first-party, 1 hands-on, 0 relaying] Watch: Prompt Engineering: GPT-6 Astra: The harness matters more than you think
- DeepSeek V4 Flash cost varies widely across nine coding harnesses - In 20 long-horizon coding tasks across nine harnesses, cost for one model varied widely, with cache reads dominating. [0 first-party, 1 hands-on, 0 relaying] Watch: Prompt Engineering: GPT-6 Astra: The harness matters more than you think
- Fish Audio S2.1 Pro voice cloning from ~15 seconds judged convincing - Bijan Bowen found instant cloning from a 15-second clip very convincing; 24 basic emotion tags plus advanced ones. [0 first-party, 1 hands-on, 0 relaying] Watch: Bijan Bowen: Fish Audio S2.1 Pro Full Test – Building A Video Translation Pipeline!
- GPT-5.6 Sol built a local transcribe-translate-dub pipeline; first render slowed speech - Using the ChatGPT Mac app at extra high, Sol set up local transcription, Qwen 3.8 Next translation and Fish Audio dubbing. [0 first-party, 1 hands-on, 0 relaying] Watch: Bijan Bowen: Fish Audio S2.1 Pro Full Test – Building A Video Translation Pipeline!
- TBC pitches neuron-derived adapters with AWS partnership and speedup claims - The Biological Computing Company claims adapters derived from living neuron models cut video generation from ~400 s/53 cents to 62 s, improve Oasis coherence and reduce hallucinations in a Cosmos 3 rollout, and shows a neuron-controlled robot. [0 first-party, 0 hands-on, 1 relaying] Watch: Amazon Web Services: In The Field – We go inside TBC.co's lab to see the future of AI effic
- GitHub talk: monolithic migration agent failed; split into orchestrator plus specialists - Speaker recounts lessons: a single agent failed, a 3-day estimate for a four-file repo, ~$100 credits burned without guardrails, autonomy sizing by reversibility, and remark that early MCP adoption cooled. [0 first-party, 0 hands-on, 1 relaying] Watch: GitHub: Jueves de Quack con Axel Labruna