We Didn’t Build the Master Algorithm. We Built an Orchestra.
This week everyone was watching GPT-6 and Opus 5.5. What made me stop was a model that doesn’t write a single word, and it sent me back to Pedro Domingos’s The Master Algorithm.
Read
Future — Data & AI.
I took the red pill.
Notes, prototypes and open field. Building systems that think with us, not for us.
This week everyone was watching GPT-6 and Opus 5.5. What made me stop was a model that doesn’t write a single word, and it sent me back to Pedro Domingos’s The Master Algorithm.
Read
OpenAI published 99.9% on ARC-AGI-3 for GPT-6 Astra. ARC Prize measured 62.7% the same day. The difference is the harness, and it explains why benchmark tables no longer compare models.
Read
Meta’s CORAL puts an LLM agent in a closed loop around a production recommender. The interesting part isn’t the numbers: it’s where they chose to plug the model in.
Read
Claude Code writes the orchestration script, launches a fleet of subagents, and the coordination doesn’t cost a single token. Third installment: from the loop to a graph that runs itself.
Read
There are around twenty specifications for giving an AI agent an identity and almost none are finished. But the serious problem is elsewhere: SCIM and SPIFFE operate on incompatible timescales, and nobody covers the point where they meet.
Read
Google published a 51-page whitepaper on the new software lifecycle with vibe coding. The conceptual framework — the spectrum, Agent = Model + Harness, the factory model — is excellent. The statistical apparatus collapses at the first question. A layer-by-layer review, with all nine headline figures traced to their primary sources.
Read
Graphify turns a repository into a deterministic knowledge graph with no embeddings and no LLM cost. Its benchmark is one of the best in the sector — and it contains the row that dismantles its own headline: four thousandths over a classic hybrid RRF. A technical read of the pipeline, the numbers and the defaults.
Read
An August paper proposes we stop installing skills into AI agents and start calling them by name. The idea is good and the diagnosis is better. The implementation has been on GitHub for three weeks, has two forks, and ships a package that doesn’t exist yet.
Read
Sixteen bacteriophages designed by a generative model infected and killed E. coli. The danger isn’t the machine that writes genomes: it’s that the lock protecting us was on a different door, and it has spent years recognising faces.
Read
Agent governance is usually discussed as doctrine. We took it to a test bench: a ~200-line gatekeeper, warm-forked microVMs, and four models unknowingly trying to break the rules. The numbers, the prompt-injection matrix, and the two leaks that taught us the most.
Read
Agent runtimes converged on the same architecture and nearly the same price. The interesting question is no longer which one to buy: it is what the one you already have actually enforces, and where the part you have to build yourself begins.
Read
Altman says we’re already inside the singularity. We separate the three creatures sharing one label: the real singularity of physics, AI’s borrowed metaphor, and science fiction’s quantum stage prop.
Read
Amodei clarifies that Anthropic never asked for open-weights models to be banned. His post is more reasonable than the headlines suggested — and even so, the real argument is elsewhere: do we want a hyper-efficient monoculture, or a seed vault?
Read
Ng’s four patterns are still the right vocabulary, but the real axis in agent architecture isn’t how many agents you have — it’s where the state lives. An honest walk from loops to graphs, and what has changed since 2024.
Read
The impression and the signal: three Claudes, a century and a half, and one shared question — when we communicate something, what exactly are we transmitting?
Read
Loop engineering without the smoke: the anatomy of an autonomous loop, how I run mine with a GitHub board and subagents, and the Claude Code parts that make it possible.
Read
One post whenever there's something worth saying. No metrics, no tricks. Just honest notes from the edge.