6 posts
ARC Prize says OpenAI's GPT-6 Astra scores 62.7% on ARC-AGI-3 with its standard harness and 99.9% with OpenAI's own memory settings. What the gap means.
Anthropic's Fable 5.1 ships with 75% cheaper cache reads, a safeguards layer your code must handle, and API changes that break history-editing harnesses.
The biggest MCP revision since launch goes stateless, formalizes extensions, hardens OAuth, and deprecates three core features. What matters in production.
Claude Code can now write its own multi-agent harness — fanning out subagents, verifying their work, and returning one answer. A sourced deep dive.
How AI-native IDEs, agentic workflows, and autonomous coding are reshaping software development.
Patterns for building reliable AI agents with tool use, memory, and multi-step reasoning capabilities.