Home / Work / Ad Library

One dataset, four surfaces.

Most companies collect data and use it once. The same competitor ad dataset now earns backlinks, captures leads, ranks for long-tail search, and answers questions inside ChatGPT and Claude.

The problem

Wilow was collecting public advertising data every day to power the product. That data sat behind a login, doing exactly one job. Meanwhile the top of the funnel was thin, the domain had no authority, and AI assistants answering questions about advertising had no reason to mention us.

The instinct in that position is to write more blog posts. The better move was to notice that the data we already held was the asset, and that it could be pointed at four different problems without collecting anything new.

What I built

Surface one, a free public tool. Anyone can search a brand and see the ads it is running, with email gates placed where the value is obvious rather than where they are annoying. It is genuinely useful whether or not you ever buy anything.

Surface two, programmatic pages. Per-brand and per-vertical pages, plus comparison and long-form guide pages, targeting the long tail rather than fighting for the head term against much larger domains.

Surface three, original research. A named, recurring index built from data already collected, covering 91 UK brands and 5,460 ads. Original statistics are the most linkable content that exists, and a named index compounds because people cite it and come back for the next edition. Every stat is attributed to the canonical page, so the authority pools in one place.

Surface four, an MCP server. The same data exposed so AI assistants can query it directly when someone asks what is working in a category. Being the tool an AI calls is worth more than being the tenth link it cites.

The research engine makes humans cite you. The MCP endpoint makes AI assistants call you. Same dataset, two entirely different distribution channels.

How it holds up

The research is deliberately constrained so it stays credible. UK-first, because that is where the reach data is reliable. The sample is named on every claim rather than implied. And nothing gets published that the data does not actually show, which rules out the revenue and spend axes that would have made the charts more exciting and the research worthless.

Operationally, brand ingestion runs on a schedule, the labelling pipeline runs in two passes with a human correction step, and every email capture point carries a privacy notice at the point of collection rather than buried in a footer.

What it produces

  • A free tool that is a lead magnet and a product demo at the same time
  • Programmatic pages competing on terms the head-term players ignore
  • A citable research asset that earns links to one canonical URL
  • An MCP endpoint that puts the brand inside AI answers, where the next generation of discovery happens

What this proves

The valuable question was never "what should we publish". It was "what do we already have that nobody else can reproduce, and how many surfaces can it serve". That is a consulting decision, and the build follows from it.

Talk about your data