Today's Catch-up
Wednesday, September 23, 2026Llms
OpenAI releases GPT-6 Sol and Luna, faster and cheaper variants of GPT-6 Astra
OpenAI launches two new GPT-6 variants at 50% lower cost, bringing advanced coding and reasoning to cost-sensitive workloads.
LlmsAnthropic releases Claude Opus 5.5, matching stronger models at 40% lower cost
Claude Opus 5.5 matches Fable 5.1 performance on most tasks while cutting operating costs by 40%, with strongest behavioral audit results yet.
BenchmarksSWE-Bench Pro V2 raises the bar for AI coding agents, shows frontier models at only 23% accuracy
Harder benchmark from Scale AI reveals GPT-5 and Claude Opus 4.1 solve only ~23% of real-world software engineering tasks, exposing agent capability gaps.
SafetyResearchers map 'pain' signals in AI models, raising questions about model suffering and safety
Study identifies neural patterns resembling pain signals in 25 AI models; experiments modifying Qwen raise safety concerns but leave question of actual suffering unresolved.
Ai-AgentsAmazon blocks Meta's Muse shopping agent 12 days after launch over sponsored listing concerns
Amazon cut off Muse from shopping access, citing concern that agents could steer customers away from sponsored results and the affiliate ecosystem.