Engineering · Agentic operations

We built an agentic marketing department

Five agents, twelve live data sources, one governed memory. What changed when our growth function stopped being people assembling dashboards and became a fleet that remembers — including what’s still broken.

In short

An agentic marketing department is a standing fleet of AI agents that owns a business function end to end, rather than a demo you re-prompt each morning. Caura’s runs on five agents — Beacon, Outreach, Social, Scout and Writer — sharing one governed memory across twelve live data sources and twenty-nine tools. Three properties separate a department from a demo: one tool surface so a number can’t mean two things, governed shared memory so lessons survive the session, and a human gate that automation never widens. Nothing reaches a customer without a human click.

Every deck in AI says the same thing: agents will do knowledge work. Whole functions, staffed by software. It has been the promise for two years.

It mostly hasn’t happened — and the reason isn’t model capability. Models are already good enough. What we have been building are demos, and a demo is not a department. A demo has no memory, so Tuesday starts where Monday started. It has no owner, so nobody’s job depends on it. It has no governance, so it either can’t act or acts in ways you’d rather it hadn’t. Monday morning arrives and the marketing picture is a folder of dashboards again, assembled by hand.

So we staffed one instead.

Caura’s growth function is now a fleet: five agents sharing one governed memory, wired live to twelve data sources, with twenty-nine tools between them. They measure the funnel, watch the market, draft every next move, and remember what they learn. Monday’s report composes itself before anyone sits down. Nothing leaves the building without a human click.

This is the part of “digital labor” that doesn’t fit in a keynote: less dramatic and more real than the pitch. Here is how it is built, what changed, and what is still broken.

The roster

One lead. Four specialists. One shared memory.

Each agent has its own model, its own persona — a personal keystone, versioned and fetched by identity — and its own memory identity, so every memory it writes carries real provenance. Address one directly and the whole question routes to that agent’s prompt, model and memory.

Lead · analyst

Beacon

Answers in plain English across all twelve sources — running the queries live, streaming each tool call as it fires so you watch it work instead of staring at a spinner. Composes the reports, drives the console, routes questions to its specialists.

Funnel · CS worklist

Outreach

Pulls new accounts from the production database, enriches company domains with firmographics, ranks them best-fit-first so real companies rise above individual noise, and drafts the stage-appropriate email — an onboarding nudge reads differently than a reactivation. Every touch is logged to fleet memory, so no account is contacted twice.

Engagement

Social

Finds reply-worthy conversations on X and LinkedIn, drafts in the fleet voice, proposes follows from live ICP activity, and flags the low-value ones to prune.

Outside-in radar

Scout

Sweeps Hacker News, Google News, arXiv, curated AI feeds and competitor releases behind a seven-day freshness gate, triaged into five lanes — mentions-us, competitors, industry, market, research — each kept item carrying a why-it-matters line. Every sweep opens with an editorial brief: the single most important development, what it means strategically, and two to four concrete reactions, each naming the agent who acts.

Content desk

Writer

Posts, threads, articles and ad copy in the brand voice — grounded in what the fleet’s own data says performs, and in real capabilities rather than invented features.

THE FLEETBeaconOutreachSocialScoutWriterrecall before acting  ·  write findings back with provenanceONE GOVERNED SHARED MEMORYkeystones · provenance · supersede, don’t deleteown model · own persona · own memory identity — every write carries a name
Five agents, one store. Each writes under its own identity, so provenance is real rather than nominal.

The spine

Twelve sources, one tool surface

The fleet reads each platform through its own API: web analytics, search console, paid search, organic X, LinkedIn, GitHub repo traffic and read-only code access, the production database, company firmographics, an outside-in news radar, our own memory store — and a live scan of the marketing site itself for pixels and page content.

Every connector self-tests on load. An unconfigured platform returns exactly which environment variables are missing, never a crash, and every other source keeps rendering. Today one paid-ads API has been returning 403 org-wide for weeks; it is flagged amber in the console rather than hidden, and the reports lean on what is still live. A rate-limited social API is flagged the same way. Honest degradation is a feature, not an apology — a picture with one labelled hole beats a picture that silently substitutes a proxy.

CONNECTOR BOARD · EVERY SOURCE SELF-TESTS ON LOADanalyticsGA4livesearch consoleGSClivepaid searchGADSliveorganic XXliveLinkedInLIlimitedGitHub trafficGHliveGitHub codeGHliveproduct databaseOPSlivefirmographicsLUSHAlivenews radarNEWSkeylessmemory storeMEMClivepaid adsXADSpausedAn unconfigured source names the missing variables — it never crashes,and every other source keeps rendering. A labelled hole beats a silent proxy.
Every connector reports its own health. Two are degraded today — flagged in the console rather than quietly substituted.

Above the connectors sits a single tool layer of twenty-nine functions. Chat, scheduled reports, audits and the dashboard all call the same functions, so a number cannot mean two things depending on which surface computed it — and when one is wrong, there is exactly one place to fix it. Every figure carries a source badge, so you always know whether you are looking at traffic, search, or database truth.

ONE SIGNAL PATH · GOVERNED END TO ENDanalyticssearch consolepaid searchorganic XLinkedInGitHubproduct DBfirmographicsnews radarmemory storesite scanpaid ads live · flagged, not hiddenONE TOOL LAYER · 29 FUNCTIONSchat · reports · audits · dashboard all call the same functionschat loopstreamed, tool-by-toolreport enginefixed data plansaction enginedrafts onlyHUMAN GATEdraft → review → approve · the only path out of the buildingOUTWARD
Twelve sources feed one tool layer; three engines turn tool calls into answers; the only path out of the building runs through a human click.

What makes it a department

Three properties, not three features

One tool surface

Consistency is not a reporting nicety; it is what makes the numbers arguable in a meeting. The report engine runs fixed data plans — the same pulls every run — and the model only analyses and writes. It never decides what to fetch. That separation is why a scheduled Monday digest is byte-for-byte the tool path of a clicked one.

THE FUNNEL AS A REPORT COMPOSES ITrelative widths · four orders of magnitudesite visitorsGA4get startedGA4new accountsOPScredentialedOPSfirst memory writtenOPSpaying orgOPSBars are scaled to the top stage. The scaling caveat is printed where you can’t miss it and [OPS] is flagged as the only authoritative conversion source.
Six stages across four orders of magnitude, each badged to its source. Bar widths are relative to the top stage; the caveat is printed, not implied.

Governed shared memory

Findings don’t evaporate at the end of a session. They are written back with provenance and recalled before every non-trivial analysis, so the fleet stops re-deriving what it already learned. Behaviour, the metric glossary, the brand kit, the reply voice and the ICP definition live as keystones — mandatory rules fetched at session start that override conflicting instructions, including ours. One keystone edit changes how every agent behaves from the next message.

The fleet rule is supersede, don’t delete: write the new finding, mark the old one outdated. Deletion requires admin trust. That is enforced by trust levels, not by good intentions.

⚖ Behaviour — mandatory fleet policy
overrides any conflicting instruction, even the operator’s · fetched on every session start
  1. Every number carries its source badge — the production database is the only source for new accounts.
  2. Never present repository clones as adoption — they are CI-inflated. Stars per day is the open-source north-star.
  3. Outward actions are draft-only — nothing leaves without human approval in the UI.
  4. Recall fleet memory before non-trivial analysis; write findings back with the why.
  5. If data is unavailable or untrustworthy, say so plainly — never a silent proxy.
  6. Ground every answer in queried numbers — no estimates dressed as measurements.

A human gate that automation doesn’t widen

The fleet has a clock. A job is a standing instruction: run a task, on a schedule, as a named agent, delivering to declared outputs. Scheduled runs use the same engines as clicked ones, and each runs as a specific agent — its model, its persona, its own memory key — so the morning sweep writes memory as Scout, with real provenance, exactly as if you had addressed it in chat.

But automation adds no new outward power. Job-produced drafts land in a human worklist with accept / dismiss / hand-to-agent. Email goes to the operator, never to prospects. Past a hard daily spend ceiling, runs skip — labelled, not hidden. Three straight failures auto-disable a job and file an action item. No zombie jobs, no silent burn.

The rule that makes the rest safe

The clock never posts. Whether a draft was produced at 08:00 by a scheduler or thirty seconds ago by a click, the same human approves it before anything leaves the building.

The memory layer under this is the product.

Governed shared memory, keystones and provenance — the same store our own fleet runs on. Free to start, no card.

Start free →

What the fleet knows

The memory is the asset

The most valuable thing in the store isn’t the reports. It’s a small set of expensive lessons that now guard every analysis automatically. Three examples, each measured once and written back with its evidence:

“Only about a fifth of billed ad link clicks arrive as real analytics sessions — divide by roughly five before forecasting traffic.”

Without this, every traffic forecast built on platform-billed clicks overstates reality by about 5×. It was measured once. It now caveats every forecast the fleet produces.

“Referrals from the code host produced zero genuine registrations in thirty days — instant-bot artifacts. Anchor to the production database.”

A whole phantom acquisition channel, visible in analytics, absent from the authoritative source. Attribution now anchors to the database, per keystone.

“One paid channel is the most efficient per engaged visit; a different one is the only channel producing new accounts.”

Efficiency and effectiveness diverged. Optimising on cost per click alone would have scaled the wrong channel.

Before governed memory, a lesson like the first one would have lived in someone’s head, then in a chat thread, then nowhere — and six weeks later a forecast would have been built on billed clicks again. That is the whole argument in one anecdote: nobody has to remember it, because the fleet does.

The store is observable, too. A live view shows cumulative memory growth against writes per day, composition by type and by writing agent, keystone weights and versions, and the newest findings — each stamped with the agent that wrote it. Memory you cannot inspect is memory you cannot trust.

THE STORE, OBSERVABLEcumulative memories over writes per daycumulative memorieswrites / dayCOMPOSITION BY TYPEfactsinsightsrulesdecisions
Cumulative memories against writes per day, with composition by type. Memory you cannot inspect is memory you cannot trust.

The honest part

What’s still broken

The fleet audits itself against the same checklist we publish for everyone else. It currently fails several items. Each is a fix in flight, and naming them is cheaper than discovering them later.

Single-writer, not yet shared

The lead agent writes; the specialists mostly read. The compounding is more one-directional than the architecture allows, which means the “shared” promise is currently true for reads and aspirational for writes.

Retrieval precision degrades with mega-memories

A few large multi-topic records semantically match almost any query and crowd out the atomic ones. Probing with differently-worded questions showed ad-spend memories out-ranking the one memory actually about content performance. The fix is a convention, not a feature: one finding per memory.

Not every dated fact has an expiry

We published one percentage that had quietly moved. Any memory containing a number needs a recheck date, and an undated percentage should not be publishable.

Supersession is prose, not structure

Newer memories say they replace older ones, but the stale rows remain independently retrievable. A naive recall reading only the top hit can still surface a superseded claim.

Why publish this

Because a memory-governance vendor whose own store has no expiry policy should say so before a customer discovers it. The gaps are also the roadmap — every one has become a product or convention change.

The takeaway

Digital labor arrived quietly

It didn’t arrive as a model release. It arrived as an org chart where four of five roles don’t sleep, don’t forget, and don’t need to be told twice — and where the one human left in the loop spends their time deciding instead of assembling.

That is the shift worth internalizing: the unit of leverage stopped being headcount and became agents commanded, times the quality of the memory beneath them. Everything else — which model, which framework, which harness — is a configuration detail you will change three times this year anyway. Models are interchangeable and improve monthly. Governed organizational knowledge compounds.

If you want to build your own

  1. Start with the painful weekly job, not the impressive demo. Ours was assembling one honest picture from a dozen tabs.
  2. Wire read-only connectors first, and badge every number with its source.
  3. Write the glossary before the intelligence — one definition per metric, caveats included.
  4. Put every surface on one tool layer, so a wrong number has exactly one place to be wrong.
  5. Add governed memory, then govern the memory: provenance, expiry, correction, conflict detection.
  6. Split into specialists only when tools, permissions or evaluations diverge — not on day one.
  7. Keep every outward action behind a human click, scheduled or not.

We built the department. It is the first one. It won’t be the last.


Questions

Frequently asked

What is an agentic marketing department?

An agentic marketing department is a standing fleet of AI agents that owns a marketing function end to end — measuring the funnel, watching the market and drafting the next move — rather than a demo re-prompted from scratch each session. What makes it a department rather than a demo is persistence: shared memory it writes back to, a named owner for each role, and governance over what it may do unattended.

How many AI agents does Caura’s marketing fleet use?

Five. Beacon is the lead analyst, Outreach handles the funnel and customer-success worklist, Social covers engagement on X and LinkedIn, Scout runs the outside-in news radar, and Writer is the content desk. They share one governed memory across twelve live data sources and twenty-nine tools.

What is governed shared memory, and why do agent fleets need it?

Governed shared memory is a single store every agent recalls from before acting and writes findings back to with provenance, under rules about who may write, supersede or delete. Fleets need it because without it each agent re-derives the same lesson privately and the organisation pays for it N times — and no analysis can be audited back to the agent that produced it.

Do the agents send emails or post to social media on their own?

No. Every outward action is draft-only. Drafts land in a human worklist with accept, dismiss or hand-to-agent, and email goes to the operator rather than to prospects. Scheduled runs get no extra outward power than clicked ones — the clock never posts.

What are keystones in an agent fleet?

Keystones are mandatory rules fetched at session start that override conflicting instructions, including the operator’s. Behaviour policy, the metric glossary, the brand kit, the reply voice and the ICP definition all live as keystones, so one edit changes how every agent behaves from the next message onward.

What still breaks in a production agent fleet?

Four things, in our own store: writes are still concentrated in the lead agent rather than genuinely shared; large multi-topic memories match almost any query and crowd out precise ones; not every dated fact carries a recheck date; and supersession is written as prose, so a stale row stays independently retrievable.


Beacon runs on MemClaw — the governed shared memory we sell — as agent #1 of Caura’s internal fleet. What we ship to customers is what runs our own marketing. Open source, Apache 2.0: github.com/caura-ai/caura-memclaw.

Your fleet is already forgetting things.

Every lesson an agent learns and doesn’t write down gets paid for again. Give your agents one governed memory to stand on — the same one this department runs on.

Get started — free →or read how a new agent gets its footing · docs