Research note 01 — current edition
The Web Context Problem
Most AI products that reason about the outside world fail for an unglamorous reason: the material they reason over is incomplete, stale, or structurally unusable. This note argues that web context is a distinct infrastructure layer with its own quality properties, and that treating it as a solved detail of prompt engineering produces systems that look correct and are not.
Aaron Grainger / 2 min / Published Nov 4, 2025
Research note 02 — current edition
The Anatomy of a Source-Linked AI Answer
A source-linked answer is not an answer with links appended. This note breaks a well-formed answer into its parts — claim, support, provenance, confidence, and refusal — and argues that the structure of the answer object matters more than the wording of the prose.
Aaron Grainger / 2 min / Published Feb 17, 2026
Research note 03 — current edition
Why Clean Content Beats Raw HTML for Most AI Tasks
Raw HTML is a rendering instruction set that happens to contain text. This note examines what is lost and gained when a page is normalized to structured Markdown, and identifies the narrow set of tasks where the markup itself is the signal worth keeping.
Aaron Grainger / 2 min / Published Jun 21, 2025
Research note 04 — current edition
The Hidden Maintenance Burden of DIY Web Data Pipelines
The first version of a web-data pipeline is usually a weekend. The cost arrives afterwards, in silent breakage, template drift, and the operational question of who notices when a source stops producing useful content. This note catalogues where that ongoing cost accumulates.
Aaron Grainger / 2 min / Published Feb 10, 2025
Research note 05 — current edition
What AI Agents Need From the Open Web
Agents interact with a web that was designed for human readers and search crawlers. This note sets out what an agent actually requires from a page, where current conventions fall short, and which of those gaps are a publisher's problem rather than an agent builder's.
Aaron Grainger / 2 min / Published May 12, 2026
Research note 06 — current edition
A Taxonomy of Web-Data Failure Modes
Debugging a web-data pipeline is easier when failures have names. This note proposes a working taxonomy across five layers — access, retrieval, normalization, extraction, and interpretation — and notes which layer each symptom usually belongs to.
Aaron Grainger / 2 min / Published Aug 19, 2025
Research note 07 — current edition
The Difference Between Finding Information and Using It
Search solves location. Most AI products fail at the step after location: turning a set of plausible pages into material a system can act on. This note separates the two problems and argues that conflating them is why 'add search' rarely fixes an unreliable product.
Aaron Grainger / 2 min / Published Dec 3, 2024
Research note 08 — current edition
From Page Retrieval to Product Reliability
Reliability in a web-enabled AI product is not an average of component accuracies; it is determined by how the system behaves when a component fails. This note traces the path from a single page fetch to a user-visible guarantee and identifies where guarantees can actually be made.
Aaron Grainger / 2 min / Published Jul 29, 2026
Research note 09 — current edition
The Practical Limits of AI Research Automation
Automated research is good at breadth and weak at judgement. This note sets out the specific judgements that resist automation — source credibility, conflicting evidence, and the difference between absence and non-existence — and suggests where a human belongs in the loop.
Aaron Grainger / 2 min / Published Apr 8, 2025
Research note 10 — current edition
A Framework for Evaluating Web-Enabled AI Workflows
Evaluating a web-enabled workflow with a single accuracy number hides where it is actually weak. This note proposes a five-dimension evaluation frame — coverage, freshness, fidelity, traceability, and recoverability — with suggested observations for each.
Aaron Grainger / 2 min / Published Aug 25, 2026