Growth Creative · Bestie AI

AI Creative Pipeline for Growth

Role

Design, working alongside engineering and growth — I own the visual side and the image-generation logic the AI follows, and designed parts of the operator UI. The multi-agent architecture itself is engineering-led.

Duration

Dec 2025 - Jul 2026

Type

AI Creative Pipeline / Tooling UX

Skills

Visual Spec for AI, Prompt Engineering, Information Architecture, Tooling UX

Tools

Figma, Claude, Cursor, Meta Ads

At Bestie, growth ads are generated by a multi-agent AI pipeline, not drawn by designers. My deliverable was no longer the interface. It was the rules the AI must follow.

A product designer by title, I stepped in when the growth team's pipeline started shipping off-brand at scale: I owned the brand-constraint layer, and contributed to the asset taxonomy and the human-in-the-loop review tools that keep AI-generated ads on-brand.

At a glance

The problems I stepped in to solve

Business problem

AI creatives drifted off-brand at generation speed: avatars misused as logos, product UI invented by the model.

My call

Stop reviewing ads one by one. Codify the brand into hard constraints the model cannot ignore, and put them in the code layer, not a hot-editable prompt.

What it changed

Both failure classes were shut off at the source. The rules ship through PR review, so no prompt edit can silently drop them.

Business problem

Captions were context-blind: the pipeline couldn't tell which asset fit which ad scenario.

My call

Treat taxonomy as design material: a 4-category use_case system so assets carry their own context.

What it changed

use_case became a required field on every asset. Captions turned context-aware, and the tag travels with each asset into every generation call.

Business problem

A pipeline that publishes to Meta without human checkpoints is a brand liability.

My call

Design the human-in-the-loop layer as a product: review queue, dry-run prompt editor, asset manager.

What it changed

Nothing reaches Meta unreviewed: every ad passes a human approve/reject queue and uploads paused by default, and a prompt edit can't even be saved without passing a dry run on a real winning ad.

If this role is new to you

It's design work you already know, pointed at a machine

What I did

I wrote invariant visual rules that every generation agent must obey.

The design skill you already know

A design system, except its only consumer is a model. Same discipline: tokens, constraints, usage rules. Different reader.

What I did

I diagnosed systematic brand drift in AI output and codified the fix into prompts.

The design skill you already know

Brand governance and design QA, running at generation speed instead of review-meeting speed.

What I did

I helped design the review queue where humans approve or reject each ad before publishing.

The design skill you already know

Human-in-the-loop checkpoint design: the approval step is a designed property of the system, not a fallback.

The system

Nine agents, one rule layer

Nine specialised agents cover the creative lifecycle: one diagnoses published ads, three generate static images (iterating winners, proposing new concepts, remaking underperformers), and five more run the video studio chain. I own the visual rule layer that governs all of them, and designed parts of the operator tools around the loop. The agent architecture itself is engineering-led.

INVARIANT VISUAL RULESbrand constraints the AI must follow, expressed inside system promptsASSETLIBRARYupload · auto-captionuse_case taxonomy ×4DIAGNOSEaudit outputagentVARIANTiterate winnersagentCONCEPTnew directionsagentREMAKErebuild losersagentREVIEWQUEUEapprove / reject→ Meta publishHUMAN-IN-THE-LOOP TOOLCHAINasset manager · PE prompt editor (dry-run) · review UI · competitor browser

The static-image path, where the rule layer does most of its work. The five video agents run a parallel chain under the same constraints.

The real creative evolution loop: generations of ad variants scored, then kept or killed

The loop the diagram abstracts: each generation of variants is scored, and the system keeps the winners and kills the rest. Figures redacted.

Before the rules

The creative was grounded in real users, not guesswork

A pipeline can only enforce rules; a human still has to decide what good creative is. Grounded in the team's user research and competitive analysis, I translated who the user actually is into the ad concepts the pipeline generates, and into what “on-brand” means for it to protect.

Who the ads speak to · de-identified archetypes

The evidence analystThe social strategistThe dating noviceThe life narratorThe pattern breakerThe creative collaborator

Before

A core user need“I feel something's off. I need proof, and a way to respond.”

After

A scenario ad conceptA tricky message on screen, three on-brand replies from the squad. “Don't know how to reply? We do.”

Every generated ad has to survive this translation: real emotional need in, on-brand creative out. That standard is exactly what the guardrails below defend.

The case · brand drift

When the machine forgets who the brand is

Left unconstrained, the pipeline drifted: character avatars were misused as logos, and product UI was invented by the model instead of being sourced from the real asset library. Each drifted ad quietly eroded brand trust, at generation speed.

The scale problem wasn't one bad ad, it was compounding: the pipeline generates on a daily automated schedule, so a bad rule reproduces its mistake every single morning until someone codifies the fix. I flagged the two failure classes and defined what “correct” looks like; the fix below shipped as a pair with engineering.

The fix

From design spec to prompt constraint

The audience of a design spec changed: it is now read by a model, not a person. I led the fix, translating visual rules into prompt-level hard constraints, turning “the avatar is never a logo” from a guideline humans agree with into an instruction the model cannot ignore (shipped with engineering).

Before

The human ruleA persona avatar is never the app logo. Chat UI is never drawn from imagination, it comes from real product screenshots.

After

The machine version (paraphrased)Hard constraints injected exactly where assets attach to the generation call; the edit instruction changed from “as shown in reference” to “COPY the provided UI VERBATIM”; and the soft escape, “if no UI asset, describe a chat scene”, was deleted outright.

Information architecture

Taxonomy as design material

I contributed to a four-category use_case taxonomy for the asset library, so the pipeline could caption and compose assets according to ad context: information architecture in service of model behaviour.

The tools

Human-in-the-loop, by design

A pipeline is only as reliable as the tools operators use to steer it. I contributed to the design of the operator toolchain: an asset manager with auto-captioning, a prompt editor with dry-run validation, an ad review queue feeding Meta publishing, and a competitor ad browser.

Operator toolchain: prompt editor, competitor browser, review queue, analytics dashboard

Impact

What the rules changed

1
The flagged drift classes stopped recurring: brand rules now live in the code layer behind PR review, so they survive every prompt iteration instead of depending on one.
2
The taxonomy became part of the team's asset contract: use_case is a required field on upload, and every generation call consumes it.
3
The checkpoints held as volume scaled: human review before publish, paused-by-default uploads to Meta, and dry-run-gated prompt edits, across a pipeline of nine specialised agents (one diagnostic, three static-image, five video).

The context these rules protected. Scaling paid growth is a whole team's work, and the numbers below are theirs, not mine — my job was to keep that volume on-brand as it grew. Over seven months monthly ad spend scaled roughly 8x while Meta CPI fell 46% and blended CPA fell 32%.

The analytics dashboard the growth work was accountable to, all figures redacted

The rules were accountable to real numbers: this is the analytics surface the growth team steered by. Values redacted.

Seven months of ad scaling: spend up roughly 8x while cost per install fell 46 percent, axis values withheld

What designers ship now

Working inside an AI-native team changed my definition of a design deliverable. Rules, taxonomies and constraints are design artifacts too, and the designer becomes a translator between design intent and model behaviour. As more teams put generation inside their creative loop, the ability to write rules a machine can follow, and design the checkpoints where humans stay in control, is becoming its own craft. This is the work I want to keep doing.

The other half of this role, designing the product itself — an AI companion's memory turned into something you can feel — is a separate story: Bestie AI · product design →