Search for AI for marketing agencies and you get two kinds of pages: lists of tools, and lists of agencies ranking themselves. Neither answers the question an agency founder is actually asking, which is closer to: what do I do about AI in my agency, this quarter, without betting the business on it.

This guide answers that question. It maps what AI can genuinely carry in an agency, the three routes to getting it (buy tools, hire, or build systems), where it pays first, what it cannot carry, and the capacity argument for building before hiring. It is written by a team that builds these systems, and it says so where that matters.

What can AI actually do for a marketing agency?

Two different things, usually conflated. Client-facing AI changes what you sell: AI-assisted creative, AI search optimisation, new service lines. Delivery AI changes how the work gets made: reports that produce themselves, briefs extracted from calls, budget checks that run nightly, first drafts that arrive for editing rather than writing.

Most of the noise is about the first kind. Most of the margin is in the second, because delivery is where agency hours actually go, and delivery work is repetitive in exactly the way models handle well. This guide is about the second kind.

One distinction worth keeping from the agency-world debate: using AI inside existing workflows is not the same as rebuilding a workflow around it. Pasting a report into a chatbot saves minutes; a system that produces the report saves the workflow. The difference compounds monthly.

The three routes: buy, hire, or build

Every option you will ever be pitched is one of these three. They are not competitors; they solve different problems, and most agencies end up with a mix. The failure mode is picking one for a problem that belongs to another.

Buy vs hire vs build: the actual decision
Buy toolsHire (people or an agency)Build systems
What you getSubscriptions: assistants, copy tools, reporting SaaSCapability in heads: an AI lead, or an outside teamCode that produces specific deliverables in your stack
Time to valueDaysMonthsWeeks per system
Cost shapePer seat or client, foreverSalary or retainer, foreverOne build each, optional upkeep
DifferentiationNone: your competitors buy the same toolsDepends entirely on who you hireYours: built around your workflows
OwnershipVendor's roadmap, vendor's termsKnowledge leaves when people doCode and credentials stay yours
Right forIndividual productivity, standard tasksStrategy, taste, client leadershipRepetitive delivery work at scale

What the mixed stack looks like in practice

A concrete shape, for a twelve-person agency: assistant subscriptions for everyone (buy, cheap, individual productivity); a built reporting system and a built pacing-and-anomaly system (the two workflows that ate the most delivery hours); and the next hire redirected from a third account manager to a senior strategist, because the account managers now carry more accounts.

Notice what did not happen: no platform migration, no AI transformation programme, no head of AI. Three decisions, each matched to the right route.

A trap in the search results

If you searched this topic, you met pages offering to be your AI marketing agency. Read them carefully: they sell AI-powered marketing to brands. If you run an agency, they are your competitors, not your solution. The service you would actually be buying from a partner like us is different: systems built into your own delivery, sold to you, owned by you.

What about no-code automation platforms?

Zapier-class and Make-class platforms sit in the buy column, and they are genuinely good at glue: a form triggers a brief, a closed deal opens a folder. Treat them as the duct tape of the stack. Where they strain is exactly where agencies need trust most: multi-source data transformation and client-facing numbers, where a broken step fails silently and nobody owns the logic. Automate notifications with them; do not ship client figures through them.

Five questions for any AI partner you evaluate

  • Who owns the code, the credentials and the output when the engagement ends? The only good answer is: you do.
  • Does the first delivery run in production on real client work, or is it a pilot beside the work?
  • What is the mechanism that stops the system inventing numbers? A reassurance is not a mechanism; an architecture is.
  • What exactly happens if we stop paying? Systems that switch off were rentals, whatever the contract called them.
  • Is the price fixed before the build starts? Open-ended AI engagements are how transformation programmes are born.

And one tell that costs nothing to check: notice the questions they ask you. A partner who proposes solutions before mapping how your delivery actually works is selling a template, whatever the deck says.

Where AI pays first in an agency

Ranked by how reliably the work repeats, which is what makes it automatable:

  • Client reporting: the clearest case, because it is monthly, structured and identical in shape across clients. It is worth its own guide: automated client reporting for agencies covers the architecture and the build-or-buy split in full.
  • Budget pacing and anomaly checks: nightly comparisons of spend against plan, and alerts when a metric breaks pattern. Small systems, outsized value, because they catch problems before clients do.
  • Brief and meeting capture: calls transcribed, decisions and actions extracted into your project tool, under human review. Kills the retyping layer between conversation and work.
  • First-draft production: outlines, ad variants, summaries. This is where tools alone genuinely work, provided a human owns the final text and the claims in it.
  • Client onboarding: access requests, tracking checks and kickoff docs generated from a standard checklist the moment a deal closes. Repetitive, high-stakes, and nobody's favourite job.
  • New-business research: prospect and competitor scans compiled into a standard format before the pitch. Lower frequency, so automate it later, not first.

The pattern across all of these: structured input, a shape that repeats every week or month, and an output a human can verify at a glance. Work with those three properties automates well; work without them (negotiation, creative direction, a crisis) does not, and pretending otherwise is how AI projects fail.

A filter before you buy or build anything

Borrowed from the better tool guides and worth applying ruthlessly: a purchase or a build must improve throughput, quality or measurability, and you should be able to say which one, with a number you will check. If you cannot, it is a distraction with a subscription fee.

What AI cannot carry

The honest list, because the failures are expensive:

  • Client-facing numbers, unsupervised. Language models are unreliable at arithmetic and recall; a report where the model writes the figures will eventually invent one. The fix is architectural: figures inserted from source APIs, the model writing only commentary. One invented number costs a client.
  • Strategy and taste. Positioning calls, creative judgment, saying no to a client: the work you charge the most for is the work AI supports least. That is an argument for automating everything else.
  • The relationship. Reports can produce themselves; trust cannot. The point of removing delivery hours is to spend them in the room with clients.
  • Your accountability. Whatever drafts the work, the agency signs it. Which is why governance deserves its own section.

The governance rules to write before you scale anything

Four rule sets, written down once, applied across every client. Skipping this is the mistake agencies regret at contract-renewal time:

  • Data rules: what client data may enter which tools, and what never leaves your accounts. Business-tier AI accounts with no-training terms are the baseline, not the answer.
  • Brand and claims rules: which statements need a source, per client. A system can enforce this mechanically; a rushed human cannot.
  • Approval rules: who signs off before anything reaches a client, and which outputs earned the right to skip review. Trust is granted per workflow, not globally.
  • Access rules: admin-controlled accounts, not personal logins, so nothing walks out with a departing employee.

The capacity argument: build before you hire

Why this became true now, and not three years ago: models crossed the threshold where they reliably extract structure from messy input and write competent commentary over verified data. The constraint moved from model capability to integration, from can AI do this to is it wired into where the work happens. Integration is a build problem, which is precisely why the delivery layer is now addressable.

Here is the thesis, stated plainly. Agency growth has one traditional answer: more clients, more account managers. Every founder knows the treadmill version: revenue grows, headcount grows faster, margin stays where it was.

AI changes the arithmetic for exactly one category of work: the repetitive, structured delivery layer. Reports, pacing checks, briefs, research packs. Hand that layer to systems, one workflow at a time, and each account manager carries more accounts at the same quality. Your team takes the next account without the next hire.

This does not make hiring wrong. It changes what you hire for: judgment, client leadership, taste, the things that cannot be built. Most agencies will keep hiring for the repetitive layer anyway, because hiring is the habit. A few will build instead. The margin difference between those two groups will not be subtle.

The maths founders actually run

Skip the vendor ROI decks and run your own arithmetic: take one account manager, list their accounts, and count the hours per month that go to reporting, pacing checks, briefs and retyping between tools. That number, times your team, times your billable rate, is the size of the prize; it is usually the salary of the hire you were about to make.

Then the second number: what those same hours produce if they move to client conversations and strategy, the work clients actually renew for. The capacity argument is not about cutting people; it is about what the same people carry.

The order of operations that works

  • Map where delivery hours actually go, per workflow, not per person. The result usually surprises founders: reporting and admin outrank creative.
  • Automate the workflow that costs the most first, in production, with its results measured on your own numbers, not a vendor's case study.
  • Only then decide the next one. Nothing continues until the last build has proven itself. Pilots that run beside the real work die quietly; systems that produce the real work stay.

Three ways this goes wrong (seen in the wild)

  • The tool graveyard: a dozen subscriptions bought in enthusiasm, each used by one person for one month. Cause: no filter, no owner, no measurement. Cost: real money and, worse, organisational cynicism about AI before the useful version arrives.
  • The eternal pilot: a demo that impresses in the all-hands and never touches real client work. Pilots are safe precisely because they do not matter; systems earn trust by producing the actual Monday deliverable with a human gate on it.
  • The unowned stack: delivery quietly rebuilt on a founder's personal accounts and a freelancer's scripts nobody else can read. It works until that person leaves. Ownership rules and admin-controlled access are boring; so are seatbelts.

A 90-day plan that does not bet the agency

  • Weeks 1-2: map delivery, workflow by workflow. Where do the hours go, who touches what, which steps are retyping. Write the four governance rule sets in the same fortnight; they cost an afternoon now and a crisis later.
  • Weeks 3-8: one build, in production, on your most expensive workflow, for real clients. Not a pilot beside the work: the actual Monday report, produced by the system, reviewed by a human.
  • Weeks 9-12: measure on your own numbers, fix what the first month exposed, and decide the second workflow. If the first build did not prove itself, stop; you risked one workflow, not the agency.

The quarter ends with either a working system and a shortlist for the next one, or a cheap, contained lesson. Both outcomes beat a year of tool subscriptions nobody measured.

How building works, if you go that route

Our version of this is the loop we run for every client: map how the work gets made, assemble the system that produces it inside your existing stack (your accounts, your credentials, nothing migrated), run it in production with every figure traced to its source, then scale to the next workflow. The deepest build we have published is the Apshan data platform, which runs the same discipline at lakehouse scale.

The entry point is deliberately small: a free 30-minute mapping call that maps where your delivery hours go and which workflow to hand to a system first. You leave with the map either way; builds start at 5,000 pounds if you want one.