←  All case studies
Client work · solo build · scope, product, code and integrations

A caption a model invents is worthless.

A content product for creators — upload a video or image, get back something post-ready. The hard part was never the writing.

My roleSole product owner & builder
Duration3 months to MVP-ready
TeamJust me
StackGoogle Video Intelligence · Claude Sonnet · FFmpeg
0
Posts ever published without human approval
4
Platforms integrated, post-only access
$0.50
Unit-cost ceiling per render and analysis
3
Subscription tiers, credits sized to the ceiling
1
Person on the whole thing
01
The problem

Fluent, on-brand, and completely forgettable.

The client wanted creators to upload a video or image and get back a post-ready asset — caption, hook, hashtags, shaped per platform — then publish to their connected accounts in one click.

Ask a language model to "write a caption for this image" and it will. It will be grammatical, cheerful, and indistinguishable from the thousand other captions generated the same way that morning. Creators don't need more words. They need words that work.

The model had no idea what actually performs. It was guessing, confidently, in a vacuum.

So the problem stopped being a writing problem and became a grounding problem.

What generic generation produced

Captions unrelated to what was actually in the frame. Hooks that didn't match the content. Hashtag sets that were plausible but wrong for the niche. All of it fluent, none of it grounded.

What it needed to produce

Copy anchored to observed performance on that platform, for content like this one — then handed to a human before it went anywhere near a real account.

02
The decision

Retrieve first. Generate second. Never the other way round.

Instead of prompting a model to invent copy, the system read what was actually in the upload, went and found what had already performed for content like it, stripped the noise out of that set, and synthesised from that evidence base. Generation anchored to observed performance rather than to model imagination.

Upload video or image Analyse location · subject background · type Retrieve top-performing captions for similar content noise stripped Synthesise caption · hook · tags per platform APPROVAL GATE The creator nothing auto-publishes Publish IG · YT FB · X Refused at upload graphic violence · sexual violence blocked before any render cost re-prompt · unlimited

The creator sits between generation and publication, not after it. Unsafe uploads never enter the pipeline at all — cheaper, and it means the system never has to decide whether to publish something it shouldn't have made.

03
The guardrails

Three calls about power, not about prompts.

01

A human before anything ships

Nothing auto-published, ever. Every caption, hook and hashtag set went to the creator for approval, with unlimited re-prompting to steer it. The model proposes; the person whose account it is decides.

02

Block at the input, not the output

Graphic violence and sexual-violence content refused at upload, before any analysis or render. Filtering output means paying to generate something you then throw away — and trusting a classifier at the worst possible moment.

03

Least privilege on account access

Post-only permissions across Instagram, YouTube, Facebook and X, with an explicit policy that the app reads nothing from a connected account. Creators hand over their reach. They shouldn't have to hand over their data too.

04
Unit economics

Cap the unit cost, not the user's freedom.

Video analysis and rendering are expensive per operation, and generous free tiers are how products like this quietly bankrupt themselves. But metering someone's creative iteration is a terrible experience — nobody wants a counter ticking down while they try to get a caption right.

$0.50
Ceiling per video render and analysis — the expensive operation, capped by design
$0.02
Ceiling per prompt exchange — cheap enough that iteration never needed rationing

So I separated the two. Unlimited re-prompting inside a generation, because that's where the value is. A monthly credit balance sized to those unit ceilings, because that's where the cost is. The constraint was the balance, never a turn limit.

Free
One generation. Enough to see whether it works for you.
Premium
Regular posting cadence, credits sized to the render ceiling.
Pro
High-volume creators, same ceilings, larger balance.
05
Evaluation

The evaluation I should have built, written out properly.

I can't retrofit this onto a project that ended. I can be specific about what it would have been — because vagueness here is exactly how "we tested it" gets said about systems nobody tested.

01

A set, not a vibe

Fifty uploads spanning the formats the client actually posts, frozen before any tuning. Every later change scored against the same fifty, so that a comparison means something.

02

Two graders, blind

Caption-to-video relevance and hook quality, scored one to five by two people who don't know which system produced which output, with the disagreement rate reported next to the score.

03

A baseline to beat

The same fifty run through plain prompting, no retrieval. If grounding doesn't beat that, the retrieval layer is cost without benefit — and I'd have had the evidence to cut it.

The gate stopped bad output reaching a creator. It never told me how often bad output was being produced.

That distinction is the one I'd want to be pressed on. A human approval gate is a guardrail — it bounds the damage. It is not evaluation — it never measures the rate. I had the first and treated it as the second, and the person absorbing the difference was the creator, one draft at a time.

06
What I'd do differently

I judged quality by eye, and I shouldn't have.

There was no systematic evaluation. Bad output — captions unrelated to the image, hooks that missed — got caught by me looking at them and by the human loop correcting them. That works at ten uploads. It does not work at ten thousand, and it gave me no way to prove the retrieval grounding was worth its cost.

Fifty uploads, blind-scored for caption-image relevance against a plain-prompting baseline, would have answered that in an afternoon. It would also have made the model-swap decision evidence-based rather than a plan I never got to run. I built the guardrails and skipped the measurement — which is the same mistake, one layer up, that I later made on my own product.

How it ended

MVP-ready at three months. The client discontinued the project before launch, so it never reached a real creator. I keep it in this portfolio anyway — the decisions were real, and most of them I'd make again.