Engineering · Agentic operations
Five agents, twelve live data sources, one governed memory. What changed when our growth function stopped being people assembling dashboards and became a fleet that remembers — including what’s still broken.
An agentic marketing department is a standing fleet of AI agents that owns a business function end to end, rather than a demo you re-prompt each morning. Caura’s runs on five agents — Beacon, Outreach, Social, Scout and Writer — sharing one governed memory across twelve live data sources and twenty-nine tools. Three properties separate a department from a demo: one tool surface so a number can’t mean two things, governed shared memory so lessons survive the session, and a human gate that automation never widens. Nothing reaches a customer without a human click.
Every deck in AI says the same thing: agents will do knowledge work. Whole functions, staffed by software. It has been the promise for two years.
It mostly hasn’t happened — and the reason isn’t model capability. Models are already good enough. What we have been building are demos, and a demo is not a department. A demo has no memory, so Tuesday starts where Monday started. It has no owner, so nobody’s job depends on it. It has no governance, so it either can’t act or acts in ways you’d rather it hadn’t. Monday morning arrives and the marketing picture is a folder of dashboards again, assembled by hand.
So we staffed one instead.
Caura’s growth function is now a fleet: five agents sharing one governed memory, wired live to twelve data sources, with twenty-nine tools between them. They measure the funnel, watch the market, draft every next move, and remember what they learn. Monday’s report composes itself before anyone sits down. Nothing leaves the building without a human click.
This is the part of “digital labor” that doesn’t fit in a keynote: less dramatic and more real than the pitch. Here is how it is built, what changed, and what is still broken.
The roster
Each agent has its own model, its own persona — a personal keystone, versioned and fetched by identity — and its own memory identity, so every memory it writes carries real provenance. Address one directly and the whole question routes to that agent’s prompt, model and memory.
Answers in plain English across all twelve sources — running the queries live, streaming each tool call as it fires so you watch it work instead of staring at a spinner. Composes the reports, drives the console, routes questions to its specialists.
Pulls new accounts from the production database, enriches company domains with firmographics, ranks them best-fit-first so real companies rise above individual noise, and drafts the stage-appropriate email — an onboarding nudge reads differently than a reactivation. Every touch is logged to fleet memory, so no account is contacted twice.
Finds reply-worthy conversations on X and LinkedIn, drafts in the fleet voice, proposes follows from live ICP activity, and flags the low-value ones to prune.
Sweeps Hacker News, Google News, arXiv, curated AI feeds and competitor releases behind a seven-day freshness gate, triaged into five lanes — mentions-us, competitors, industry, market, research — each kept item carrying a why-it-matters line. Every sweep opens with an editorial brief: the single most important development, what it means strategically, and two to four concrete reactions, each naming the agent who acts.
Posts, threads, articles and ad copy in the brand voice — grounded in what the fleet’s own data says performs, and in real capabilities rather than invented features.
The spine
The fleet reads each platform through its own API: web analytics, search console, paid search, organic X, LinkedIn, GitHub repo traffic and read-only code access, the production database, company firmographics, an outside-in news radar, our own memory store — and a live scan of the marketing site itself for pixels and page content.
Every connector self-tests on load. An unconfigured platform returns exactly which environment variables are missing, never a crash, and every other source keeps rendering. Today one paid-ads API has been returning 403 org-wide for weeks; it is flagged amber in the console rather than hidden, and the reports lean on what is still live. A rate-limited social API is flagged the same way. Honest degradation is a feature, not an apology — a picture with one labelled hole beats a picture that silently substitutes a proxy.
Above the connectors sits a single tool layer of twenty-nine functions. Chat, scheduled reports, audits and the dashboard all call the same functions, so a number cannot mean two things depending on which surface computed it — and when one is wrong, there is exactly one place to fix it. Every figure carries a source badge, so you always know whether you are looking at traffic, search, or database truth.
What makes it a department
Consistency is not a reporting nicety; it is what makes the numbers arguable in a meeting. The report engine runs fixed data plans — the same pulls every run — and the model only analyses and writes. It never decides what to fetch. That separation is why a scheduled Monday digest is byte-for-byte the tool path of a clicked one.
Findings don’t evaporate at the end of a session. They are written back with provenance and recalled before every non-trivial analysis, so the fleet stops re-deriving what it already learned. Behaviour, the metric glossary, the brand kit, the reply voice and the ICP definition live as keystones — mandatory rules fetched at session start that override conflicting instructions, including ours. One keystone edit changes how every agent behaves from the next message.
The fleet rule is supersede, don’t delete: write the new finding, mark the old one outdated. Deletion requires admin trust. That is enforced by trust levels, not by good intentions.
The fleet has a clock. A job is a standing instruction: run a task, on a schedule, as a named agent, delivering to declared outputs. Scheduled runs use the same engines as clicked ones, and each runs as a specific agent — its model, its persona, its own memory key — so the morning sweep writes memory as Scout, with real provenance, exactly as if you had addressed it in chat.
But automation adds no new outward power. Job-produced drafts land in a human worklist with accept / dismiss / hand-to-agent. Email goes to the operator, never to prospects. Past a hard daily spend ceiling, runs skip — labelled, not hidden. Three straight failures auto-disable a job and file an action item. No zombie jobs, no silent burn.
The clock never posts. Whether a draft was produced at 08:00 by a scheduler or thirty seconds ago by a click, the same human approves it before anything leaves the building.
Governed shared memory, keystones and provenance — the same store our own fleet runs on. Free to start, no card.
What the fleet knows
The most valuable thing in the store isn’t the reports. It’s a small set of expensive lessons that now guard every analysis automatically. Three examples, each measured once and written back with its evidence:
“Only about a fifth of billed ad link clicks arrive as real analytics sessions — divide by roughly five before forecasting traffic.”
Without this, every traffic forecast built on platform-billed clicks overstates reality by about 5×. It was measured once. It now caveats every forecast the fleet produces.
“Referrals from the code host produced zero genuine registrations in thirty days — instant-bot artifacts. Anchor to the production database.”
A whole phantom acquisition channel, visible in analytics, absent from the authoritative source. Attribution now anchors to the database, per keystone.
“One paid channel is the most efficient per engaged visit; a different one is the only channel producing new accounts.”
Efficiency and effectiveness diverged. Optimising on cost per click alone would have scaled the wrong channel.
Before governed memory, a lesson like the first one would have lived in someone’s head, then in a chat thread, then nowhere — and six weeks later a forecast would have been built on billed clicks again. That is the whole argument in one anecdote: nobody has to remember it, because the fleet does.
The store is observable, too. A live view shows cumulative memory growth against writes per day, composition by type and by writing agent, keystone weights and versions, and the newest findings — each stamped with the agent that wrote it. Memory you cannot inspect is memory you cannot trust.
The honest part
The fleet audits itself against the same checklist we publish for everyone else. It currently fails several items. Each is a fix in flight, and naming them is cheaper than discovering them later.
The lead agent writes; the specialists mostly read. The compounding is more one-directional than the architecture allows, which means the “shared” promise is currently true for reads and aspirational for writes.
A few large multi-topic records semantically match almost any query and crowd out the atomic ones. Probing with differently-worded questions showed ad-spend memories out-ranking the one memory actually about content performance. The fix is a convention, not a feature: one finding per memory.
We published one percentage that had quietly moved. Any memory containing a number needs a recheck date, and an undated percentage should not be publishable.
Newer memories say they replace older ones, but the stale rows remain independently retrievable. A naive recall reading only the top hit can still surface a superseded claim.
Because a memory-governance vendor whose own store has no expiry policy should say so before a customer discovers it. The gaps are also the roadmap — every one has become a product or convention change.
The takeaway
It didn’t arrive as a model release. It arrived as an org chart where four of five roles don’t sleep, don’t forget, and don’t need to be told twice — and where the one human left in the loop spends their time deciding instead of assembling.
That is the shift worth internalizing: the unit of leverage stopped being headcount and became agents commanded, times the quality of the memory beneath them. Everything else — which model, which framework, which harness — is a configuration detail you will change three times this year anyway. Models are interchangeable and improve monthly. Governed organizational knowledge compounds.
We built the department. It is the first one. It won’t be the last.
Questions
An agentic marketing department is a standing fleet of AI agents that owns a marketing function end to end — measuring the funnel, watching the market and drafting the next move — rather than a demo re-prompted from scratch each session. What makes it a department rather than a demo is persistence: shared memory it writes back to, a named owner for each role, and governance over what it may do unattended.
Five. Beacon is the lead analyst, Outreach handles the funnel and customer-success worklist, Social covers engagement on X and LinkedIn, Scout runs the outside-in news radar, and Writer is the content desk. They share one governed memory across twelve live data sources and twenty-nine tools.
Governed shared memory is a single store every agent recalls from before acting and writes findings back to with provenance, under rules about who may write, supersede or delete. Fleets need it because without it each agent re-derives the same lesson privately and the organisation pays for it N times — and no analysis can be audited back to the agent that produced it.
No. Every outward action is draft-only. Drafts land in a human worklist with accept, dismiss or hand-to-agent, and email goes to the operator rather than to prospects. Scheduled runs get no extra outward power than clicked ones — the clock never posts.
Keystones are mandatory rules fetched at session start that override conflicting instructions, including the operator’s. Behaviour policy, the metric glossary, the brand kit, the reply voice and the ICP definition all live as keystones, so one edit changes how every agent behaves from the next message onward.
Four things, in our own store: writes are still concentrated in the lead agent rather than genuinely shared; large multi-topic memories match almost any query and crowd out precise ones; not every dated fact carries a recheck date; and supersession is written as prose, so a stale row stays independently retrievable.
Beacon runs on MemClaw — the governed shared memory we sell — as agent #1 of Caura’s internal fleet. What we ship to customers is what runs our own marketing. Open source, Apache 2.0: github.com/caura-ai/caura-memclaw.
Every lesson an agent learns and doesn’t write down gets paid for again. Give your agents one governed memory to stand on — the same one this department runs on.
Get started — free →or read how a new agent gets its footing · docs