The source ad
Rank 1 of the top-15 set pulled from Facebook page 763671557066427 — the Sweat app (Kayla Itsines’ platform), fronted by one of its Pilates program creators.
The whole ad is one person, one take, one location, and phrase-synced captions. No edit, no product footage, no logo card. Set: a bright, high-key living room — blown-out windows, sheer curtains, tropical greenery outside, white sofa. Wardrobe: hot-pink sports bra (matches the caption accent color) over taupe leggings. Camera at eye level, waist-up framing, dead center.
Full transcript (ElevenLabs, single speaker, 0.1–20.7s): “My Pilates workouts only cost $2.60 per week on an annual plan. I literally can’t think of anything you can get for $2.60 these days. So if you sign up to Sweat, you’ll get my at-home Pilates program, a weekly workout planner, high protein recipes, water tracking, and so much more. So sign up and start improving your strength, flexibility, and mobility right now.”
Frame-by-frame breakdown
Sampled at 2 fps across all 20.8s. The ad is four beats; captions are white pill overlays at chest height, one phrase at a time, with prices and keywords colored Sweat pink.
“My Pilates workouts only cost $2.60 per week on an annual plan.”
Warm smile, hands clasped, then an open-palm flourish on the price. Captions: My Pilates workouts → only cost $2.60 per week → on an annual plan. A specific, absurdly small number lands inside the first 1.5s — before any scroll decision. “On an annual plan” pre-frames the commitment honestly, so the offer survives the landing page.
“I literally can’t think of anything you can get for $2.60 these days.”
Open-hands shrug on “literally,” laughing delivery. The price is repeated — twice in eight seconds — and anchored against everyday inflation (“these days”). This is the relatability beat: she’s not selling yet, she’s agreeing with the viewer that everything is expensive.
“So if you sign up to Sweat, you’ll get my at-home Pilates program, a weekly workout planner, high protein recipes, water tracking, and so much more.”
The longest beat — five list items, each with its own caption pill, counted off on her fingers: at-home Pilates program a weekly workout planner high protein recipes water tracking and so much more. The finger-enumeration gesture is the visual for this beat; it keeps a static shot alive for 7.5 seconds.
“So sign up and start improving your strength, flexibility, and mobility right now.”
Prayer-clasp on “sign up,” open-hand rolls on the benefit trio, small emphatic fist on “right now.” Ends mid-energy — no logo card, no end frame. Captions: So sign up and start → improving your strength → flexibility and mobility → right now.
Caption system
- Placement: horizontally centered, ~48% height (chest level, never covering the face).
- Style: white rounded-rect pill, dark charcoal text, sentence case, ~1–5 words per pill, subtle drop shadow.
- Emphasis: prices and offer keywords in Sweat pink ; everything else neutral.
- Timing: pills swap exactly on phrase boundaries — effectively word-timed karaoke at phrase granularity.
Why this ranks #1
The formula, extracted — this is what we’re actually replicating, not the pixels.
- Price-specificity hook. “$2.60 per week” outperforms “affordable” because a precise odd number reads as computed truth, and the weekly unit makes an annual subscription feel like pocket change.
- Zero production distance. One take, no cuts, no polish — it reads as a person talking, not an ad, which buys the first 3 seconds on feed.
- Repetition of the number. The price is said twice and shown twice in pink before the offer is even named.
- Enumerated value stack. Five concrete items against a tiny price — the classic “stack until the price feels wrong” direct-response move, made visual with finger-counting.
- Benefit trio close. Strength / flexibility / mobility — outcomes, not features, in the last breath before the CTA button below the video.
- Captions carry 100% of meaning. Fully watchable on mute; the pink highlights do the selling even at zero volume.
The 50Queen adaptation
Same four-beat skeleton, re-grounded in our audience (US women 48–65), our offer (12-week programme, founding price $49, one payment), and our brand voice (healthy, steady, independent — never toned, burn, transform).
Adapted script v1 (~21s, matched to source pacing)
00–04 · “My 12-week programme for women over fifty works out to about four dollars a week.”
04–08 · “One payment, no subscription — honestly, what else can you get for four dollars these days?”
08–16 · “When you join 50Queen, you’ll get three short workouts a week, all at home, no equipment — just a chair — with my voice coaching you through every session, and so much more.”
16–21 · “So join today, and start feeling strong, steady, and independent — right now.”
For TTS, spell the brand as “Fifty Queen” so the voice model reads it correctly. Caption pills keep the source’s white-pill system but swap pink for mulberry on “$4 a week”, “just a chair”, and the benefit trio.
Set, wardrobe, delivery
- Set: keep the source’s high-key bright living room — sheer curtains, greenery through windows, white sofa edge. It codes “home workout” and flatters skin tones.
- Wardrobe: Roxie’s sage gymwear from content/_character/roxie-charsheet-sage-gymwear — the locked ad-work identity sheet. Prompt the garment, not the body (per the character README’s moderation lessons).
- Delivery: warm, smiling, continuous small gestures — clasp, open palms, finger-count on the stack, prayer-clasp on the CTA. The gestures are load-bearing: they’re what keeps a 21s single shot watchable.
Production pipeline
This is a different problem from the rank06 ChillFit remake: that was silent movement; this is a lip-synced talking head. Three candidate pipelines, raced on Beat 1 only — exactly the w1-pilot playbook that already worked.
Option A — Start-frame + audio-driven talking avatar
- Still: Nano Banana Pro start frame — refs: sage-gymwear charsheet + a source keyframe for framing/set. Approve 1 of ~4 candidates (as in w1-pilot).
- Voice: generate the four beats with Roxie’s locked TTS preset (Qwen Audio 3.0 TTS Flash, preset f6448975) — the same voice that already narrates the workout content, so ad and product sound identical.
- Video: Higgsfield audio-driven talking-avatar generation (“speak” / UGC-presenter flow) per beat, driven by the approved still + beat audio. Lips come from the audio directly.
- Assembly: stitch the four beats with alternating 2–3% punch-ins at beat boundaries — the standard UGC soft-cut, which also hides any identity flicker between clips.
Why it’s the favorite on paper: lip accuracy is the make-or-break of this format, and audio-driven generation solves it structurally instead of hoping. Identity is pinned by the approved still. Main risk is gesture liveliness — speak models can produce a stiff talking bust.
Option B — H3 motion-transfer + lip-sync pass (rank06-proven base)
- MiniMax H3 per beat: --video-references = source beat segment, --image-references = charsheet plus the approved still (the “best of both” fix noted in the w1-pilot README — charsheet alone drifted).
- Then a lip-sync pass over the H3 output using the Roxie VO.
- Wins: inherits the source’s exact gesture performance — the finger-counting and palm flourishes that make the ad. Risks: two-pass cost, lip-sync fighting the transferred mouth motion, and the identity drift already observed in w1.
Option C — Start frame + script-as-prompt (Eugene’s proposal)
Skip the dedicated lip-sync machinery entirely: one video-generation call per beat, with --start-image = the approved Roxie still (establishes character and set) and the script line plus gesture direction written into the text prompt — e.g. “she speaks warmly to camera, saying: ‘My 12-week programme for women over fifty works out to about four dollars a week’ — hands clasped, then an open-palm flourish on the price.” Model: MiniMax H3, per the standing model preference (Seedance reads too polished).
- C1 — native-dialogue model: if the model generates its own speech audio (Veo-3-class), lips are accurate but the voice is the model’s, not Roxie’s locked TTS preset — breaking voice continuity with the product’s coaching audio. Redubbing with Roxie’s voice over foreign lip timing reintroduces the mismatch we were avoiding.
- C2 — silent model, dub after (recommended sub-variant): prompt the performance, let the model produce plausible talking-mouth motion, then lay Roxie’s TTS underneath. Lips flap approximately rather than precisely — but at feed size (360px wide, captions carrying the meaning, most views muted) approximate may genuinely be good enough.
Wins: the simplest and cheapest pipeline — one generation per beat, no speak-model dependency, no second pass; identity and set pinned by the start frame (the w1-pilot’s S2 arm confirmed start-image locks identity and set). Risks: the same w1 evidence cuts the other way — S2’s prompted movement softened the source’s gesture energy; text prompts steer gestures but can’t choreograph them, so phrase-synced moves (finger-counting the value stack in Beat 3) are luck, not control. Timing is also loose: the model paces the line itself, so the VO must be fitted in post (trim, micro-stretch) instead of driving the clip’s length.
Beat-1 comparison — A vs B vs C
| Dimension | A · Still + speak model | B · Motion transfer + lipsync | C · Start frame + prompt |
|---|---|---|---|
| Lip accuracy | Structurally exact — lips generated from the audio | Post-hoc pass fighting the transferred mouth motion | Approximate flap (C2) or exact-but-wrong-voice (C1) |
| Voice identity | Roxie TTS drives generation | Roxie TTS drives the lipsync pass | Roxie TTS dubbed after (C2); lost (C1) |
| Gesture quality | Model-dependent; speak models trend stiff talking-bust | Inherited from source performance — best in class | Prompt-steered: natural but unchoreographed, high take-to-take variance |
| Identity stability | Pinned by approved still | Drift observed in w1 (charsheet-only refs) | Pinned by start frame (w1 S2 evidence) |
| Timing control | Exact — audio length drives clip | Exact — source segment drives clip | Loose — model paces itself; VO fitted in post |
| Pipeline complexity | Medium (still → TTS → speak) | High (two generation passes) | Lowest (one call per beat) |
| Est. cost / beat | speak rate TBD | ~24 + lipsync | ~24 |
| Fails if… | Gestures read dead | Lips read dubbed / face drifts | Lip-flap reads fake at full attention |
The honest framing: C is the baseline every fancier option must beat. If C2’s approximate lips pass the eye test in-feed, its cost and simplicity win outright — A and B only earn their extra machinery if the dub mismatch is visible enough to hurt trust in a talking-head format whose entire premise is “a real person telling you a price.”
Caption & finish pass
- Transcribe our own final VO (ElevenLabs, word timestamps) and cut caption pills exactly on phrase boundaries — same sync mechanism the source uses.
- Pill overlays via the rank06 ffmpeg recipe — including the logged gotcha: single-frame PNG inputs need -loop 1 -t N or fades stay invisible.
- Master at 1080×1920; deliver 9:16 plus a 4:5 crop (the ads-copy fix-plan showed placement crops bite later if skipped).
- Light music bed, ducked well under VO; source is essentially VO-only, so this is optional polish.
Cost & risks
Rates from the w1-pilot ledger: stills 2 cr, 6s H3 clip 24 cr. Talking-avatar (“speak”) pricing needs a check against current Higgsfield rates before the race.
| Stage | Scope | Est. credits |
|---|---|---|
| Start-frame factory | 4 Nano Banana Pro candidates | ~8 |
| VO | 4 beats, Roxie TTS preset | ~4 |
| Beat-1 race | Option A + B + C clips (C adds one H3 call) | ~75–105 |
| Full build (winner) | Beats 2–4 (+1 retry margin) | ~90–150 |
| Overlay/finish | ffmpeg only | 0 |
| Total | ~175–265 (less if C wins — full build drops to ~72–96) |
| Risk | Mitigation |
|---|---|
| Lip-sync uncanny valley | That’s what the Beat-1 race decides; kill the format if neither arm passes an honest eye test. |
| Identity drift across beats | Approved still in every generation’s refs; punch-in cuts mask residual flicker; Soul ID remains the documented fallback if drift hits 2+ beats. |
| Moderation on gymwear prompts | Sage set is modest; describe the garment, not the body (character README rules). |
| 21s static shot feels dead with a stiff avatar | Punch-ins per beat + gesture-forward prompting; Option B exists precisely because its gestures are inherited from the source. |
| Claim accuracy | “$4 a week” must match the live founding49 ladder at launch time; re-verify web/src/lib/plans.ts before publishing, and keep “about four dollars” hedged in VO. |
Note on the source material: Option B uses Sweat’s footage as a motion reference during generation. Fine for an explorative spike; nothing of theirs ships in the final asset either way — structure and formula aren’t copyrightable, their pixels and their creator’s likeness are off-limits.
Next steps
- 1 · Eugene: approve script v1 (or pick variant B’s hook), and green-light the three-arm Beat-1 race budget (~85–120 cr incl. stills + VO). No generation runs until this go-ahead.
- 2 · Verify Higgsfield’s current talking-avatar product + pricing; lock the Option A tool choice.
- 3 · Run the start-frame factory → Eugene approves one still.
- 4 · Beat-1 race → side-by-side vs source → verdict → batch the winner.
- 5 · Caption/finish pass, 9:16 + 4:5 masters, then decide whether it enters the 50Queen campaign as a variant against “Says Who 15s.”