Posted by Alumni from Substack
September 6, 2026
The AI industry has developed a peculiar new benchmark: can you finish reading a model's system card before its replacement ships' This week, OpenAI, Anthropic, Meta, and Google turned the release calendar into a competitive sport. Somewhere, an engineer is still updating last week's model comparison spreadsheet. Please give them space. OpenAI's GPT-6 Astra arrives with an expansive pitch spanning computer use, software engineering, and scientific work. OpenAI reports 98% on FrontierMath Tier 4 alongside improvements in computer interaction. The practical ambition is clear: models that navigate software, execute complicated workflows, and deliver usable work with less supervision. OpenAI is also updating the Codex harness to accelerate computer use, underscoring how much performance depends on the combination of model and surrounding infrastructure. The benchmark chart is becoming a job description'and the software around the model is becoming part of the resume. Anthropic's Claude... learn more