Skip to content

Search

ESC
All tags

Posts

Stylised cover with a small ARC-AGI-3 game board of coloured blocks on the left and a scatter plot on the right where most dots sit below a diagonal line labelled same actions as humans, with the caption 62.7% standard, 99.9% provider adapter, not AGI
24 min

GPT-6 Astra vs ARC-AGI-3: One Model, Two Very Different Scores

ARC Prize says OpenAI's GPT-6 Astra scores 62.7% on ARC-AGI-3 with its standard harness and 99.9% with OpenAI's own memory settings. What the gap means.

Read
Diagram of one glowing model core splitting into two labelled deployments: Fable 5.1 behind a shield marked safeguards on and generally available, and Mythos 5.1 behind a padlock marked invite only, Project Glasswing
30 min

Claude Fable 5.1: One Model, Two Names, Three Breaking Changes

Anthropic's Fable 5.1 ships with 75% cheaper cache reads, a safeguards layer your code must handle, and API changes that break history-editing harnesses.

Read
Diagram of stateless MCP: self-contained requests flowing through a load balancer to interchangeable server instances, with a session box crossed out
6 min

MCP Grows Up: What's Changing in the 2026-07-28 Spec

The biggest MCP revision since launch goes stateless, formalizes extensions, hardens OAuth, and deprecates three core features. What matters in production.

Read
Dynamic workflows diagram: a prompt fans out to parallel agent() calls, passes an adversarial verify gate, then converges to a single synthesized result
23 min

Inside Claude Code's Dynamic Workflows

Claude Code can now write its own multi-agent harness — fanning out subagents, verifying their work, and returning one answer. A sourced deep dive.

Read
Futuristic developer workspace with AI assistants
5 min

The Next Generation of Developer Tools

How AI-native IDEs, agentic workflows, and autonomous coding are reshaping software development.

Read
Agent architecture flowchart diagram
4 min

Designing AI Agent Architectures

Patterns for building reliable AI agents with tool use, memory, and multi-step reasoning capabilities.

Read