For those who build it
The Craft
The practitioner's complete, concept-level playbook of AI today — novelties included.
An agent is a model plus a harness. Everything hard lives in the harness.
010203040506070809101112
Agents vs Workflows
The foundational distinction — and the discipline of choosing.
The Five Patterns
Chaining, routing, parallelisation, orchestrator-workers, evaluator-optimiser.
Harness Engineering
Guides and sensors — the new senior discipline of the field.
Memory & Context
The four memory types and the four operations of context engineering.
Tools & MCP
Tool design principles and the standard interface for exposing them.
RAG & Its Evolutions
From naive top-k to agentic, graph, and self-correcting retrieval.
Multi-Agent, Handoffs & A2A
Topologies, transfer contracts, and the inter-agent wire protocol.
Steering
Mechanistic activation control and operational instruction-fidelity.
Code & Doc Indexing
Indexed vs runtime exploration — and the cost of phantom APIs.
Evals & Observability
Three evaluation surfaces. Traces with bodies. The issue lifecycle.
AFK & Autonomous Agents
The Ralph loop, agent fleets, and the HITL→on-rails→AFK gradient.
The Failure Taxonomy
A field guide — memory, reflection, planning, action, system faults.
Worked example — Evals
Pillar III · The Craft
Build the eval surface
Three evaluation surfaces — unit evals on discrete steps, regression suites, and continuous production trace sampling.
Pillar II · The Operating Model
Standardise the eval pipeline
The platform team provides shared eval infrastructure, dashboards, and quality gates that every agent must pass.
Pillar I · The Groundwork
Set the quality bar
Leadership defines acceptable accuracy, latency, and cost thresholds per risk tier, and ties them to the governance framework.