Tuesday, September 8, 2026
Coverage: 60 videos reviewed (0 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.
New today
OpenAI said its internal system produced a Lean-verified Navier-Stokes blowup proof
OpenAI said on Sept. 7, 2026, according to presenter Fahd Mirza, that an internal AI system produced a Lean-verified proof that fluid can form a singularity, using about 10,000 agents that exchanged nearly 5 million messages over roughly 88 hours. A message shown on screen states existence of forced blowup in R3 and T3. NYU's Tristan Buckmaster said in an X thread, as read by the presenter, that OpenAI research lead Sebastian Bubeck told him an internal model had produced a roughly 100-page proof by the same narrow approach as his own work with an Anthropic-employed co-author. Buckmaster said he was offered a joint release or a write-up crediting the model. Buckmaster also said he asked whether the model had been trained on or had access to a private Codex session where he drafted the work, was told the model did not look up user data, and received no answer on training. The presenter relayed the claims; no independent verification appears in the video.
- Evidence: 0 first-party, 0 hands-on, 1 relaying
- Watch: Fahd Mirza: OpenAI AI Solves Navier-Stokes ... By Stealing a Mathematician's Work?
Creators report GPT-6 Astra in Codex built games, apps and edited video in single sessions
Several creators reported hands-on results with OpenAI's GPT-6 Astra, mostly through Codex, in videos uploaded Sept. 8, 2026. Julian Goldie said Astra with Blender MCP built a racing game in 7 minutes, later laggy, with a tunnel rebuild in about two minutes, and that it fixed Hermes Desktop voice mode in about 45 seconds from a screenshot. Nate Herk said Codex with Astra and Hyperframes cut a 65-second intro to 28 seconds in two iterations (about 18 and 10 minutes). Two Minute Papers' host said Astra wrote a ray tracer and reproduced a honey-coiling paper simulation as single-page HTML files, the latter in under an hour. The AI Advantage showed a city-management app said to be built by Astra without prompt or cost details. These are single-run, self-reported results. Goldie also said a month earlier he would have chosen Claude but now prefers ChatGPT/Codex, and noted the island game controls moved in only one direction. Julian Goldie's yi4Al__H5fU#0 repeats content from his livestream WwcLgc6gwj0.
- Evidence: 0 first-party, 3 hands-on, 1 relaying
- Watch: Two Minute Papers: GPT-6 Astra Changes Everything (high hype); Nate Herk: GPT-6 Astra Finally Solves AI Video Editing (full guide)
Tencent released Hy4 Preview, a 770B-parameter open-weight model, per sponsored video
Tencent released Hy4 Preview, an open-weight mixture-of-experts model under Apache 2.0, according to the presenter of a sponsored AI Code King video. He said it has 770B total and 49B active parameters, a context over 1 million tokens, and weights on Hugging Face in BF16 and FP8. The presenter relayed vendor benchmarks of 92.3 on GPQA Diamond, 85.4 on Terminal Bench and 82.9 on SWE-bench multilingual, plus a Tencent internal blind evaluation (163 experts, 203 engineering tasks) scoring it 2.99 out of 4 against 2.94 for Kimi K3 and 2.92 for GLM 5.3. He said it trails Opus 5 and GPT 5.6 on most tasks. Tencent said the model helped optimize its own training pipeline, raising end-to-end throughput about 31.8% against its baseline. The presenter said the weights are about 1.8 TB in BF16 and need about 900 GB in FP8. Figures are vendor claims relayed secondhand.
- Evidence: 0 first-party, 0 hands-on, 1 relaying
- Watch: AI Code King: Hy-4 Preview (Fully FREE): RIP Astra? This model is PRETTY CRAZY! (high hype)
OpenAI launched GPT Image 2.5 Sunburst and Flare in API, ChatGPT and Codex
OpenAI launched GPT Image 2.5 in two variants, Sunburst and Flare, available in the API, ChatGPT and Codex, according to its launch video. OpenAI described Sunburst as its most capable image model, with sharper detail, lighting and textures and better edit consistency, and Flare as the fastest. OpenAI said Flare is over 50% faster than GPT Image 2 at equal quality, a vendor claim without measurement shown. Both support transparent backgrounds.
- Evidence: 1 first-party, 0 hands-on, 0 relaying
- Watch: OpenAI: Introducing GPT-Image-2.5 in the API
Continuing stories
Also notable
- Speakers report GPT-6 Astra usage draws heavily on limits but is included in subscriptions - Julian Goldie said one 15-to-20-minute Astra session used about 15% of his weekly usage on the Pro plan (usage reset that day, next reset the 15th). [0 first-party, 1 hands-on, 1 relaying] Watch: Julian Goldie: GPT- 6 Astra: Blender + Website Design + SEO
- Two Minute Papers host summarizes Astra paper: safer, but monitorability decreased - Host of Two Minute Papers said the 117-page GPT-6 Astra paper reports that, at higher reasoning effort, the model is less successful at evading monitoring of its thoughts and is safer than predecessors. [0 first-party, 0 hands-on, 1 relaying] Watch: Two Minute Papers: GPT-6 Astra Changes Everything (high hype)
- Kokotajlo says OpenAI announced added security and AI monitors after an internal-AI incident - On Machine Learning Street Talk, Daniel Kokotajlo said OpenAI announced the previous day it would improve security and have other AIs monitor new models in training and evaluations, with a human notified within 0.5 hour of a suspected hack. [0 first-party, 0 hands-on, 1 relaying] Watch: Machine Learning Street Talk: How Many Narrow AIs Could Behave Like One Superintelligence - Daniel K
- Google announced Gemini 3.8 flash, Lyra 3.5 and WeatherNext 3, per secondhand recaps - A Julian Goldie clip said Google announced Gemini 3.8 flash (tool use, self-checking, about a 1 million-token context), the Lyra 3.5 music model in the Gemini app, agentic video understanding using up to 88% fewer tokens, and WeatherNext 3 forecasts in 5-km blocks updated hourly. [0 first-party, 0 hands-on, 1 relaying] Watch: Julian Goldie: Google Just Dropped CRAZY AI Updates! 🤯 (high hype)
- Domyn said its Domyn Large reasoning model, derived from a 355B model, is on Azure Foundry - Domyn said Domyn Large, a finance-oriented reasoning model with extended context and thinking, is available on Azure Foundry and is not open weights. [1 first-party, 0 hands-on, 0 relaying] Watch: NVIDIA Developer: Specializing AI for Regulated Industries - How Domyn Uses NVIDIA Nemot
- Domyn released Domyn Small, a 10B open-weight reasoning model under the MIT license - Domyn released Domyn Small, a 10B open-weight reasoning model under the MIT license, with weights on Hugging Face and Foundry. [1 first-party, 0 hands-on, 0 relaying] Watch: NVIDIA Developer: Specializing AI for Regulated Industries - How Domyn Uses NVIDIA Nemot
- Domyn said it joined an EU consortium to train a 400B+ model over 12 months from Sept. 1 - Domyn said it joined the Europa consortium with Fraunhofer under the European Commission's Frontier AI challenge. [1 first-party, 0 hands-on, 0 relaying] Watch: NVIDIA Developer: Specializing AI for Regulated Industries - How Domyn Uses NVIDIA Nemot
- AWS Strands and Anthropic described Model Hardware Standard, an MCP-like standard for robot control - An AWS Developers speaker described the Model Hardware Standard (MHS), which the speaker said would standardize robot entry points the way MCP did for tools. [1 first-party, 0 hands-on, 0 relaying] Watch: AWS Developers: AWS Partners with Anthropic on the Model Hardware Standard
- Microsoft launched an AI gateway tier of Azure API Management in preview - Microsoft launched an AI gateway tier of Azure API Management in preview, for publishing and governing models and tools (MCP servers, OpenAPI, connectors), according to a Microsoft Reactor session. [1 first-party, 0 hands-on, 0 relaying] Watch: Microsoft Reactor: Secure and Govern Agents, MCP Servers, and AI Tools with AI Gateway
- Mastra host says Conductor moved workspaces to the cloud, added multiplayer, API and iPhone app - Conductor's CEO said on Mastra's video that workspaces now run on a remote computer, so agents keep running when the app quits. [1 first-party, 0 hands-on, 0 relaying] Watch: Mastra: Multiplayer Coding Agents in the Cloud - Charlie Holtz, Conductor
Models & learning
- Presenter says Hy4 Preview completed game and expense-audit tasks in Tencent WorkBuddy - In a sponsored video, the presenter ran four tasks with Hy4 Preview in Tencent's WorkBuddy desktop agent app, one run each: a canvas platformer, a three.js racing game, a 24-claim expense audit and a 10-slide deck. [0 first-party, 1 hands-on, 0 relaying] Watch: AI Code King: Hy-4 Preview (Fully FREE): RIP Astra? This model is PRETTY CRAZY! (high hype)
- Presenter reports Codex backtest of WeatherNext 3 on Kalshi NYC weather markets showed unproven edge - All About AI's presenter said he used GPT-6 Astra via Codex at high reasoning to write a backtest, trading bot and AWS deployment for Kalshi NYC weather markets. [0 first-party, 1 hands-on, 0 relaying] Watch: All About AI: Building a GPT-6 Kalshi AI Trading Bot From Scratch (full guide)
- AI Engineer workshop measures vLLM about 15x a Hugging Face baseline on Mistral 7B - In an AI Engineer workshop on an H100 with Mistral 7B, presenters measured a Hugging Face baseline of about 51 tokens per second and said default vLLM gave almost 15x that throughput; prefix caching increased it further. [0 first-party, 1 hands-on, 0 relaying] Watch: AI Engineer: Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay S
- AI Engineer workshop puts Mistral 7B KV cache at about 131 KB per token - A presenter said the KV cache on Mistral 7B costs about 131 KB per token, half a GB at 4K context, and that 80 users at 4K is about 42 GB, which overflows a 24 GB GPU. [0 first-party, 1 hands-on, 0 relaying] Watch: AI Engineer: Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay S
- Fahd Mirza tests MiniCPM5-2B Q8 GGUF: 5.3 GB VRAM, mixed quality, not recommended - In a llama.cpp test on an Ubuntu GPU machine, Fahd Mirza measured the MiniCPM5-2B Q8 GGUF using 5.3 GB of GPU memory including KV cache, which he said could drop to about 2 GB without it. [0 first-party, 1 hands-on, 0 relaying] Watch: Fahd Mirza: MiniCPM5-2B GGUF: Runs Great, Until It Doesn't