AI Catchup — Saturday, September 5, 2026
Saturday, September 5, 2026Agi
OpenAI's GPT-6 Astra hits 99.9% on ARC-AGI-3, Brockman declares AGI reached
OpenAI's latest model achieves near-perfect scores on advanced reasoning benchmarks, prompting leadership to claim AGI milestone.
ReasoningAnthropic's Claude formally proves Fermat's Last Theorem in 11 days
Claude autonomously completed a multi-week mathematical proof that stumped mathematicians for centuries, showcasing novel reasoning capabilities.
CodingNVIDIA's Nemotron beats the best human at the Coding Olympics
NVIDIA's model outperformed the top human competitor in the Coding Olympics, setting new benchmark for coding AI capabilities.
InferencePerplexity's ROSE serving stack beats vLLM on speed and latency
Perplexity open-sources ROSE, an inference optimization stack that outperforms industry-standard vLLM in serving speed and response latency.
AgentsOnly 2.6% of AI agent tools actually work reliably, Cohere research reveals
Cohere's ATE dataset shows that while thousands of tools exist for AI agents, fewer than 3% function reliably—exposing a critical reliability gap.