How we made OUR iPhone 18 AD THAT WENT VIRAL

Jeroen De SmetJeroen De SmetCo-founder, Callbuddy · Sep 16, 2026
·9 min read
How we made OUR iPhone 18 AD THAT WENT VIRAL

We made an ad for Callbuddy for the launch of our iPhone 18 Pro + Pro Max cases. It went viral. And one question kept coming back, in the comments and in our DMs, more times than we could answer one by one: how did you make this? So here it is. The full workflow, every step, every tool, including the parts that went wrong and the things we'd do differently next time. Credit where it's due: the entire video was built by Matthias, who is the genius behind every shot in it. What follows is his workflow, written up more or less the way he ran it. AI opens an enormous number of doors. It is also a lot harder to steer than it looks from the outside. If this saves you some of the time it cost us, it was worth writing up.

The workflow

  1. Write the story
  2. Generate a storyboard
  3. Create the visual building blocks
  4. Save everything as reusable Elements
  5. Translate the script into video prompts
  6. Generate scene by scene and correct as you go
  7. Edit
  8. Music and sound design
  9. Assemble the final video

1. Write the story

We started with the story. We worked out the broad concept with ChatGPT first. Only once that felt right did we develop it into a detailed script, covering not just what happens, but the dialogue, the visual jokes, the small reactions from the characters, and the transitions between scenes. We think this is the step that matters most. AI video gets impressive at the shot level very quickly, but without a real story you end up with a collection of nice, unconnected images. Before generating anything, we wanted to be able to answer: What happens in each scene? What does the viewer need to understand at that moment? Where is the humour? Where does the pace need to go up, and where should it slow down? Which information has to be visually clear? How do we move logically from one scene to the next? The more of this you lock down up front, the less time you spend mid-generation trying to work out what your video is supposed to be.

2. Generate a storyboard

Next we tried to give the story visual shape in Higgsfield. Our first approach was the Higgsfield Popcorn app, which generates storyboards automatically. We had ChatGPT break the full script into small storyboard prompts and fed those into Popcorn. That worked well for checking quickly whether the broad strokes of the story made sense visually. We would do this differently next time. We'd run step 3 first and create the main characters, locations and props. Then we'd build the storyboard ourselves from those images using the image generator in Higgsfield, for example with Nano Banana Pro. There are two big advantages to that order. You get far more visual consistency from the storyboard onwards, and the images you make for the storyboard can be reused later as reference images for your video generations. You're not burning credits on images that serve no purpose afterwards. The storyboard doesn't need to be perfect. It mainly has to answer one question: does the story work visually?

3. Create the visual building blocks

Then we used Nano Banana Pro in Higgsfield to create every visual building block that would recur in the video. In our case that included: The agent, in different outfits and states. The other characters, some of them in multiple outfits. Every important location, from Apple Park to our own Callbuddy office. Props: the iPhones, the Callbuddy cases, the cars, the CAPTCHA screen. This is probably the single most important step in the whole workflow, because these images define what your world looks like. It matters even more for a product video. An AI model cannot be allowed to invent a slightly different case, phone, outfit or office in every shot. If a character appears ten times, you want to have decided in advance what that character looks like. Same for locations and products. The more you define up front, the fewer decisions the video generator has to make on its own, and the less freedom it has on the elements that matter, the more consistent the end result. We'd rather over-prepare here than under-prepare. Make several reference images of a character if you know you'll need them in different conditions later: normal outfit, damaged outfit, wet hair, a different jacket, a specific expression. It saves an enormous amount of time down the line.

4. Save everything as reusable Elements

We then saved those images in Higgsfield Cinema Studio as Characters, Locations and Props, each with its own tag. That meant that while writing a video prompt we could point precisely at: The right character. A specific location. The exact iPhone. The correct Callbuddy case. A particular car or prop. Instead of re-describing what everything looks like at every generation, you build a small visual library first. It makes generating much more manageable, and it raises the odds that separate shots will actually sit together in the edit.

5. Translate the script into video prompts

Next we had ChatGPT split the full script into individual video prompts. Our main advice here: be as specific as you can. Don't leave anything important to the video generator's interpretation. So the prompt didn't just describe what happens in the scene, but how we wanted it filmed. For example: a wide shot of the character walking down a corridor. Then a close-up of her face. She looks at a door. Quick cut to her hand trying the handle. You're effectively describing a mini edit. Alongside that, we specified things like camera angle, camera movement, shot size, time of day, weather, lighting, pace, actions, character reactions, where someone is looking, sound effects, and any dialogue. Pace is worth describing explicitly. During the infiltration we wanted slower, furtive movement. During the escape, fast multishots and a lot of motion. For certain comedic beats we deliberately wanted a longer silence, or a shot that holds a fraction too long. Dialogue and sound effects went straight into the prompts. Music we deliberately left out, and we'd say so explicitly in every prompt: generate no music. If every individual AI clip arrives with its own music, you either have to strip it out afterwards or you end up with several pieces of music layered over each other. Adding music only at the end lets you build one consistent soundtrack across the whole video. It saves a lot of work later.

6. Generate scene by scene and correct as you go

We generated the video with Higgsfield Cinema Studio and Seedance 2.5. We deliberately didn't try to generate the whole film in one go. We worked scene by scene, often shot by shot. The shorter the generation, the more control you usually have. You can try a longer sequence when the action is relatively simple. If the model understands that movement well, it saves time. But for complex action, splitting everything up worked better for us. After each generation we checked a few things: Did the character's face stay consistent? Is the movement right? Is the phone still correct? Has the Callbuddy case changed? Is it clear what the character is looking at? Is the direction of movement right? Does the last frame sit logically next to the following shot? As soon as something went odd, we changed the prompt or the reference image and generated again. One lesson worth passing on: don't always try to fix a bad result by adding more text to the same prompt. Sometimes the action is simply too complex for one generation. Split it instead. Rather than one clip of someone walking to a door, opening it, looking inside, recoiling and running off, make four short shots. It gives you far more control in the edit.

7. Edit

For the edit we went with CapCut this time. First we downloaded every usable generation from Higgsfield and grouped them into folders per scene in CapCut. That's where the real puzzle starts. From one generation you might use the first two seconds. From another, only the middle. A third clip might have perfect movement but hallucinate at the end. So don't think of a generation as a finished video you have to use in full. Treat it as source material. You're looking through all of it for the fragments that hold up visually and that together tell the story. In the edit you can then hide a lot of problems by cutting at the right moment. A hand starting to look wrong? Cut before it happens. A face that shifts after three seconds? Use the first two. An action that doesn't fully land? End shot A just before it and start shot B just after. AI video doesn't only get made during generation. A large part of it comes together in the edit.

8. Music and sound design

For the music we used ElevenLabs. We described the kind of soundtrack we needed and how it should evolve across the video. You can be quite specific: start quiet and tense, build slowly, accelerate during the chase, then deliberately drop away for a comedic beat. That's useful, because it lets the music follow the structure of your video instead of running underneath it. For additional sound effects we used Envato among others. Footsteps, doors, cars, clicks, interfaces, room tone, and the small effects that give actions impact. Sound design ends up making an enormous difference. A generated video with weak audio still reads clearly as AI. As soon as the movements carry the right sound, everything feels far more believable.

9. Assemble the final video

Then it all comes together. The best fragments. The right timing. Dialogue. Sound effects. Music. Any text or titles. And above all, a lot of small cuts. For us that's the essence of this workflow: lock the story and your visual world down first, then build your characters, locations and products, then generate short scenes very deliberately, and use the edit to turn all of it into one consistent film. The biggest mistake would be to jump straight to the video generator and expect AI to come up with the film for you. The more decisions you make yourself up front, the better AI works as a production tool. Generating is only one part of it. The real work is in the thinking, the preparation, the selecting and the editing. That's the only reason we could make a video in which nothing was really filmed, and have it still feel like one piece. The core idea: AI video works best when you don't let AI decide everything. The stronger your story, reference images and shot planning are up front, the more consistent and usable your generations become.

Written by

Jeroen De SmetJeroen De SmetCo-founder, Callbuddy

Keep reading