Field notes · AI Agency

    AI Video Generation for Marketing: Remotion-Powered Pipelines That Convert.

    How to build an AI video generation pipeline for marketing using Remotion, TTS-first scene timing, nano-banana-2 character consistency, and programmatic motion graphics that actually convert.

    8 sections
    AI Agency
    10
    AI Video Generation for Marketing: Remotion-Powered Pipelines That Convert

    AI video generation for marketing works when you stop treating it like magic and start treating it like a pipeline. The winning stack in 2026 is Remotion for programmatic rendering, a TTS-first scene sync model where audio drives timing instead of the other way around, nano-banana-2 for character consistency across shots, and a motion graphics layer that turns flat AI footage into something a brand would actually pay for. Built right, one operator can ship 50+ branded videos a week.

    Short answer: A production-grade AI video pipeline has four layers. Script and voice are generated first (TTS-first). Scenes are timed to the audio waveform, not the other way around. Character and brand consistency comes from a referenced image model like nano-banana-2 that locks faces, products, and palette across shots. The whole thing renders through Remotion (React-based programmatic video) so motion graphics, captions, and overlays are deterministic. The result: a marketing video in under 5 minutes, repeatable across thousands of variations.

    Why AI video is finally usable for marketing

    Two years ago, AI video meant Sora demos that looked impressive but couldn't render the same character twice. The pipelines built around those models produced one-off art pieces, not marketing assets. A brand cannot ship a hero video where the spokesperson's face changes between scenes. A coach cannot run a 30-day content calendar where their AI avatar drifts every clip.

    What changed in 2025-2026: three things stacked. Reference-locked image models (nano-banana-2 and similar) made character consistency a solved problem. TTS quality (ElevenLabs v3, OpenAI's voice models) crossed the line where audiences stop noticing it isn't human. And Remotion matured as the deterministic rendering layer that lets you compose all of this with code instead of clicking around in After Effects.

    The combination is what makes marketing-grade AI video work. Each piece is mature on its own. Wired together, you get a pipeline that ships variations as fast as you can write briefs.

    The TTS-first principle: audio drives timing, not the other way around

    Most beginner pipelines do this backwards. They generate visuals first, then try to time voiceover to match. That produces awkward cuts, scene durations that don't fit what's being said, and lip-sync problems that get worse the longer the video runs.

    TTS-first inverts the order. You generate the voiceover before any visual exists. Then you parse the audio for word-level timestamps (most modern TTS APIs return these natively, or you can run Whisper on the output). Each sentence and each keyword now has an exact start and end time. The visuals are scheduled against those timestamps.

    The practical workflow:

    1. Script in JSON, not prose. Each chunk is one logical scene with a sentence of voiceover, a visual prompt, an on-screen text cue, and a list of B-roll keywords.
    2. Render the voiceover first. Stitch all chunks into one audio file. Capture word-level timestamps during generation.
    3. Lock scene durations to the audio. A scene's start time equals the timestamp of its first word; its end time equals the timestamp of the next scene's first word minus 100ms for breathing room.
    4. Generate visuals against those durations. Now your image or video model knows it needs a 4.2-second clip, not a 4.0-second one.

    TTS-first pipeline is a video generation architecture where text-to-speech audio is produced before any visual asset. Word-level timestamps from the audio drive every downstream timing decision: scene length, caption appearance, B-roll cuts, motion graphic triggers. This eliminates the most common failure mode in AI video, which is visuals that don't align with what's being said.

    Character consistency via nano-banana-2 (and how to use it)

    nano-banana-2 is Google's reference-conditioned image model. You feed it one or more reference images (a face, a product, a brand palette) and a prompt, and it generates new images that keep those references locked across outputs. For marketing video, this is the unlock that lets you run a recurring character (a founder avatar, a mascot, a spokesperson) across every clip without face drift.

    The way to wire it into a video pipeline:

    • Build a reference set once per character. 3-5 high-quality images: front, three-quarter, profile, with consistent lighting. Store these in your asset library. This is the character's permanent identity.
    • Generate scene images conditioned on the reference set. For each scene in your script JSON, the prompt is the visual description from your script. The reference images are passed alongside. Output: a still image of your character in that scene's situation.
    • Animate the still. Pass the generated still to a video model (Veo 3, Kling, Runway Gen-4) with a short motion prompt. Most of the time you don't need cinematic motion; subtle parallax, head tilt, or camera push is enough.
    • Cut on TTS timestamps. Remotion stitches the resulting clips back to back, gated by the audio timeline you built earlier.

    The same approach works for products. A reference set of your client's SKU lets every scene render that exact bottle, that exact phone case, that exact pair of shoes. Without reference conditioning, you get a different product every shot and the video is unusable.

    Production benchmark from our experience running this pipeline: a 30-second branded UGC-style video, scripted to rendered MP4, takes 3-5 minutes of wall-clock time on consumer-tier API plans. Cost per video lands between $0.40 and $1.20 depending on how much video model time is needed. A human-edited equivalent in CapCut or After Effects takes 45-90 minutes per video at $30-80 of contractor time. The economics flip hard at any volume above 5 videos per week.

    Remotion as the render layer (and why it beats FFmpeg-only)

    Remotion is a React framework for programmatic video. You describe the video as a tree of components (Scene, Title, Caption, Logo, BackgroundVideo) and Remotion renders to MP4. Every prop is data-driven, so you can produce 1,000 variations of the same video by looping over a list of inputs.

    People sometimes ask why not just FFmpeg. FFmpeg is the rendering engine underneath, but writing FFmpeg filter graphs for anything complex (animated captions, branded lower-thirds, scene transitions, multi-track motion) gets miserable fast. Remotion gives you React components and CSS animations. Animating an element on screen at second 4.2 is ``, not a 12-line filter complex.

    What you build in Remotion for a marketing pipeline:

    • A Scene component that takes audio start time, duration, video asset, and caption text as props. Internally it positions the video, fades in/out at the right frames, and renders the caption with brand styling.
    • A Captions component that reads the TTS word-level timestamps and renders TikTok-style word-by-word highlighting. This single piece alone is responsible for most of the conversion lift on short-form video.
    • A Brand component that applies the client's logo, font stack, and color palette across every scene. Swap the brand config object and the same video re-renders for a different client.
    • Motion graphics layer: animated counters, arrow overlays, sound-bite zooms, screen-record annotations. These are React components with Remotion's `interpolate` and `spring` functions for animation.

    The whole thing renders headless on a cloud worker. Remotion Lambda is the path of least resistance: you deploy the project to AWS Lambda, hit an endpoint with your scene data, get an MP4 URL back. Render times are roughly 30-60 seconds for a 30-second 1080p video.

    The full pipeline, end to end

    Here is what a single run looks like when an operator (or an AI agent) ships one video:

    1. Brief in. Topic, target audience, desired length, CTA. Either typed by a human or generated by an upstream LLM from a content calendar.
    2. Script generation. An LLM expands the brief into structured JSON: an array of scene objects with voiceover sentence, visual prompt, caption text, B-roll tags, and motion graphic cues.
    3. Voiceover render. ElevenLabs or equivalent TTS generates the full audio in one pass for consistent prosody. Word-level timestamps come back with the audio.
    4. Scene timing pass. Code aligns each script scene to the audio timestamps and computes durations.
    5. Image generation per scene. nano-banana-2 generates the still for each scene, conditioned on the character/product reference set.
    6. Video animation per scene. Each still is passed to a video model with a motion prompt. Output: short MP4 clips, one per scene.
    7. Remotion render. Scene metadata, audio file, video clips, captions, brand config are passed to a Remotion Lambda render. Output: final MP4.
    8. Distribution. The MP4 is uploaded to the client's scheduler (Make, Buffer, native platform APIs) and queued for posting.

    Steps 1-7 are fully automated. A human reviewer can drop in at step 7 (preview the final cut, approve or reject) or skip review entirely once the pipeline is dialed in for a given client. The marginal cost of video N+1 is API tokens plus render seconds.

    Where this actually converts in marketing

    Not every video format benefits equally from this pipeline. The formats that work best share two properties: high volume requirement, and tolerance for AI aesthetics (or absorption of AI aesthetics into the brand voice).

    • Short-form social (TikTok, Reels, Shorts). 15-60 second educational, hook-driven, or UGC-style clips. The format expects fast cuts and word-by-word captions. AI-generated B-roll fits naturally because the audience is scrolling, not analyzing. This is the highest-volume use case and the one where the pipeline returns the most leverage.
    • Personalized outbound video. Sales videos that name the prospect, reference their company, and pitch a tailored angle. nano-banana-2 keeps your face consistent; the TTS speaks their name and details from a CRM record. One template, infinite variations.
    • Product explainer carousels. Short looping videos for ad creative, Amazon listings, or landing pages. Reference-locked product imagery solves the consistency problem that killed earlier attempts at AI ad creative.
    • Newsletter and blog companion videos. Turn long-form content into 60-second video summaries automatically. Same script, different voice, different brand wrapper, ships across 5 clients on the same day.

    The formats where this pipeline still struggles: anything requiring real human presence and trust (testimonials, founder talking-head content for high-ticket sales) and anything where viewers will pause and zoom in (real estate listings, B2B case studies with specific data on screen). Use the pipeline where it shines; keep the camera for the rest.

    Use this pipeline when: you need to ship 10+ videos per week, the audience scrolls fast, the format is short-form, and consistency across episodes matters more than cinematic polish.

    Skip this pipeline when: the video is a single hero asset for a brand launch, viewers will scrutinize details on screen, or the human-trust factor is core to the conversion (sales letters to enterprise buyers, founder vlogs aimed at high-intent audiences).

    Common pitfalls when building this

    Most teams break the pipeline in one of three predictable ways.

    They generate visuals first. Then they fight the timing for the rest of the project. TTS-first is non-negotiable; if you're tempted to skip it because the voiceover step feels like a delay, you're going to spend ten times that delay fixing scene durations later.

    They skip the reference set and try to prompt their way to consistency. Prompt-only consistency works for one generation in three; the other two will have face drift, product drift, or palette drift that ships an off-brand video. Spend the 30 minutes to build a proper reference set per character and per product.

    They over-engineer the motion graphics layer before the core pipeline is stable. Captions, brand colors, and basic scene transitions are 80% of the perceived quality lift. Don't build animated counters and screen-record annotations until the boring stuff renders cleanly.

    Frequently asked questions

    What is the difference between AI video generation and AI video editing?

    AI video generation creates new footage from text or image prompts using models like Veo, Kling, or Runway. AI video editing takes existing footage and modifies it (cuts, captions, color, retiming) using tools like Descript or CapCut's AI features. The pipeline described here is generation-led with editing-style assembly via Remotion, so it's not a substitute for editing your own footage when you have real video to work with.

    Can I run this pipeline without writing code?

    Pieces of it, yes. Tools like HeyGen, Synthesia, and Pictory expose hosted versions of subsets of this stack. The tradeoff is you give up control over brand styling, render economics, and the ability to scale to thousands of variations cheaply. If you're producing under 5 videos per week for yourself, a hosted tool is fine. If you're an agency producing for clients at volume, a code-based pipeline pays for itself within the first month.

    How long does it take to build this pipeline from scratch?

    A working v1 with TTS, image generation, video animation, and Remotion rendering is roughly 2-3 weeks of focused dev work for someone comfortable with React and serverless. Most of that time is in the script-to-scene JSON contract, the timestamp-alignment code, and the Remotion components for captions and brand styling. The AI model calls themselves are 20 lines each. Production-hardening (error handling, retries, queue management, client workspace isolation) takes another 2-4 weeks.

    Is nano-banana-2 the only option for character consistency?

    No. The pattern of reference-conditioned image generation is supported by several models: Flux Kontext, Ideogram Character, Midjourney's character reference, and others. nano-banana-2 currently has the strongest reference adherence in our testing for faces and product shots, but the architecture of your pipeline shouldn't lock you to one model. Treat the image generation step as a swappable module so you can switch when something better ships.

    How do I keep AI video from looking obviously AI?

    Three things help disproportionately. First, keep clips short (2-4 seconds each) so motion artifacts have no time to accumulate. Second, lean into formats where AI aesthetics are already accepted (animated explainer, motion-graphics-forward, UGC-style with overlay captions covering most of the screen). Third, layer real audio: human-recorded voiceover when you have it, real ambient music, real sound effects. The audio doing the heavy lifting buys the visuals enormous slack.

    Can this pipeline produce ads that comply with Meta and TikTok ad policies?

    Yes, as long as you handle disclosure where required. Meta and TikTok both require AI-generated content involving people or political topics to be labeled. The platforms have built-in toggles for this in their ad managers. The pipeline output is just an MP4; compliance happens at the upload step. For non-political, non-impersonation content, AI-generated ad creative runs without issue on both platforms today.