Sunday, September 6, 2026
Coverage: 34 videos reviewed (2 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.
New today
Continuing stories
- OpenAI releases GPT-6 Astra; creators and OpenAI-affiliated users report strong agentic results - OpenAI reportedly released GPT-6 Astra to paid ChatGPT plans, the API and AWS, emphasizing long-running computer use. [1 first-party, 2 hands-on, 3 relaying] Watch: How I AI: GPT-6 Astra made YouTube thumbnails on the first try
- Head-to-head: Astra won 10 of 15 tasks over Fable 5.1, cheaper overall but slower - Nate Herk scored Astra 10 wins of 15 use cases: Fable total 9h35m and $513.36 vs Astra 11h19m and $326.98. [0 first-party, 1 hands-on, 0 relaying] Watch: Nate Herk: I Tested GPT-6 Astra vs Fable 5.1 on 15 Real Use Cases
- OpenAI rates Astra 'critical' for cyber capability, restricts access and reports jailbreak refusal stats - Per a secondhand video, Astra is the first OpenAI model rated critical for cyber; claims 100% on ExploitBench, two zero-days found, refusal of 91.5% of disallowed cyber requests vs 59% for GPT-5.6. [0 first-party, 0 hands-on, 1 relaying] Watch: Julian Goldie: OpenAI Astra Just Crossed a Dangerous AI Threshold (high hype)
- Nvidia reportedly acquires Hugging Face for about $12.9B, pledging it stays open - Two channels say Nvidia is buying/confirmed buying Hugging Face for just under $13B; one cites 18M developers and ~$150M annualized revenue and Jensen Huang saying it will remain open. [0 first-party, 0 hands-on, 2 relaying] Watch: Sam Witteveen: NVIDIA Doubles Down on Local AI With PAIR
Also notable
- OpenAI reportedly paused parts of Astra training for two weeks after Hugging Face-related incident - Video claims OpenAI paused parts of training for 2 weeks to tighten security and monitoring, restarting its largest RL run on August 28. [0 first-party, 0 hands-on, 2 relaying] Watch: Julian Goldie: OpenAI Astra Just Crossed a Dangerous AI Threshold (high hype)
- Panelist cites Nvidia quarterly revenue about $96B, up 106% year on year - Cited from the prior week's earnings; speaker hedges the exact figure and uses it to argue AI demand is not a bubble. [0 first-party, 0 hands-on, 1 relaying]
- DeepSeek V4 Flash on four RTX Pro 6000s: 33 tok/s single agent, 364 tok/s at 16 agents - Ziskind measured FP4-expert/FP8-attention DeepSeek V4 Flash at 33, 62, 116, 364 tok/s for 1, 2, 4, 16 agents then dropping. [0 first-party, 1 hands-on, 0 relaying] Watch: Alex Ziskind: All That VRAM Needs a Bigger Brain
- Open-weight models: GLM 5.3 released, Qwen 3.8 Flash Next reportedly tops Opus Max index score, local speed reports - Presenter says GLM 5.3 is open weights; Qwen 3.8 Flash Next reportedly surpassed Opus Max's 55 on the Artificial Analysis index; Qwen 3.8 27B tuned to ~300 tok/s (380 batched). [0 first-party, 1 hands-on, 3 relaying] Watch: Sam Witteveen: NVIDIA Doubles Down on Local AI With PAIR
- NVIDIA SkillSpector scans agent skills for injection and exfiltration; static mode missed natural-language injection - Open-source scanner gives 0-100 risk scores via regex/AST/YARA plus optional LLM pass. [0 first-party, 1 hands-on, 0 relaying] Watch: Fahd Mirza: How to Scan AI Agent Skills for Hidden Malware: NVIDIA SkillSpector
- Google launches Lyria 3.5 music model across Gemini app, API, AI Studio, Flow and Vids - Presenters say Lyria 3.5 has Clip (30s) and Pro (full song) API models, better vocals and picture-to-song input. [0 first-party, 0 hands-on, 1 relaying] Watch: Julian Goldie: Google AI Studio + Lyria 3.5 Is CRAZY! (high hype)
- Gemini 3.8 Flash described as Google's smartest Flash with think/tool/check loop and ~1M context - Secondhand description; can use tools, check work and read about a million tokens. [0 first-party, 0 hands-on, 1 relaying] Watch: Julian Goldie: Google Gemini NEW Updates are WILD! (high hype)
Models & learning
- NVIDIA releases PAIR v0.1, Apache-2.0 router spreading local agent requests across machines - PAIR proxies Ollama and LM Studio ports and distributes requests such as sub-agent calls across home-network machines; Windows, Linux, Mac; early 0.1. [0 first-party, 0 hands-on, 1 relaying] Watch: Sam Witteveen: NVIDIA Doubles Down on Local AI With PAIR
- Spark-X2.5 4B open model: Apache 2, 1M context claim, but slow overthinking and weak long-tail translation - Model card claims 3:1 sliding/full attention, 1M context, 200+ languages, ~20T tokens. [0 first-party, 1 hands-on, 0 relaying] Watch: Fahd Mirza: Spark X2.5 4B: What a 4B Model Can and Can't Do Locally
- BS bench reportedly finds newer frontier models worse at detecting nonsense premises - Reproducible benchmark reportedly shows newer models incl. [0 first-party, 0 hands-on, 1 relaying] Watch: GitHub: Are AI code reviews getting worse?
- KV caching lifts Qwen3 0.6B from ~4 to ~27 tok/s on Mac mini M4 MPS - Raschka's measurements; also CPU 5 to 29 tok/s with cache. [0 first-party, 1 hands-on, 0 relaying] Watch: Sebastian Raschka: Build A Reasoning Model Scratch 2: Loading a Base Model, Text Generati