
Gemini image (gemini-2.5-flash-image) · tools/video_stills.py · a few cents
WTF?Everything the film went through, in the order it happened, shown as the things themselves, with what each step cost.
Made with AI. Not slop, though: fifty-one renders of the edit, thirty-nine rewrites of the piano, ninety commits. This page is that work.
Every shot was drawn first with an image model, a few cents a frame, and laid on a timeline before any footage was bought. Draft one had four rooms and a man sitting on the pad under the desk. From draft two on, one frame was the reference and every other frame was generated from it — same camera, same furniture. Even then the bookshelf slid across the room five more times, until the rule: a frame is laid beside its neighbours on a contact sheet and checked item by item before anyone sees it.
WTF?










The stills in sequence, each held for its scripted seconds, silent. It cost nothing and settled the timing — “feels good for a first version” — so the clips could be bought against a cut that already worked.
Silent by default, so the words are cards. The first design was white type on slate, with a black band at the foot; it read like an obituary. The cards moved onto the walks' own colour fields, edge to edge, with the app's type — and that decision came back the next day, when the whole palette changed.



The first video model was Google's Veo: eight seconds from a still and a prompt, €2.85 a clip. The walk came out well. The pad waking up did not: in four side-on takes the belt ran the wrong way, a cover slid over it, and once it peeled up like a mat. Shooting it from above with a printed mark on the belt finally moved it the right way — and looked “ugly and technical”. Then the wall: ten generations a day. Nine of the ten clips never made the film.


Kling 3.0 Pro, through a subscription composer first and by API later. The same first frame, side-on: the belt moved the right way on the first try. Every live shot was redone.



Every version of the edit, as a filmstrip. The first cuts swap Veo takes in and out; v6 is the first with Kling; v8 and v9 take the notes from watching v7; v12 is the one that became v1.0. Click a strip to see it large.





Her screen during the meeting is not footage; it is composed from the app's real menu-bar window over nine generated faces, then animated so the faces move, with the countdown swapped in frame by frame — 0:03, 0:02, 0:01, “Let's walk.”




“The same person, same room, at different times of the day, doing different things while walking — a fast-paced cut.” Nine scenarios were drawn, checked on sheets, then generated: a call, a movie, a smoothie, a stretch, a book, a phone. Then sitting at the desk between walks. Then standing. The first four clips cut together were enough: “that supercut is amazing.”







Fifty-four seconds, silent: the finding, the pad waking on its own, the call with the walk taken along, the walk, the supercut, the card. Everything after this is refinement.
A piano piece that starts calm and grows exuberant under the second supercut. Six takes from a music-generation API came back as one texture each — nothing changed at the moments the film changes. Splicing two together was “a mess”. So the piece was written as notes instead: a motif, a bass, a run built from the motif, played by a sampled grand piano. The first written draft threw “random arps and chords at it”; the second was “totally chaotic — go back to reading the scores and start over”; the third went into the film as v2.0, and was then rebuilt some thirty times, one change per render, until the notes fell on the cuts. Each take below with the note it got.
I love Philip Glass. Could you figure out a way to use an API to create a song, maybe a piano sonata ‘New York’ that starts calm and becomes more exuberant during the second supercut?
Every take came back as one texture from start to end — the plan asked for a change at 12 s and at 34 s, and none had one.
Take 5 does NOT have any change at 12s or at 34s — what are you trying to say here. The piece IS NOT EVOLVING but I asked you several times at this point.
Whoa! This is a mess. What did you just do. No. You can't just splice together and increase volume. It should be ONE piece. Maybe you need to simply compose it for real in its entirety and then find a good way to turn those notes into an audio file without AI-gen.
SO MUCH BETTER! It's still a bit wonky, especially in the transition from part 2 to 3, and overall the theme isn't there yet — BUT this is a commit and an update of our video online. Consider this v2.0.
The score, pass by pass — every render that was committed, oldest first. The third is the one called v2.0; the last is the one in the film.
With the API in hand at sixty-seven cents a clip, the room was lent to other people: an older man, an older woman, a student, a man in his middle years — walking, then sitting — and two-person scenes. They joined the supercut, and the cut was put on the beat: every edit a multiple of half a second, two beats a cut in the first round, one in the second.


All thirty-four clips of the room, by what happens in them: fourteen walking, twelve sitting, five standing, three together — six seconds each, twenty-six of them in the film's supercut. Then the supercut itself, every version it went through, and the opener the same way.
Seen all together for the first time — the app, the site, the film — every colour field was a night. A PDF first: the fields as they were, then three looks. Golden hour won. Each walk kept the hues of its own night palette and only the light changed; the cards were re-rendered, the film re-cut. It also cost a detour: the re-render was first done at the tools' defaults, which lost the crossfades and doubled the fields' drift, and had to be redone from the exact recipe of the previous cut.





Two days, two people, one of them a model. The money went on generations; the tools that did the rest — the piano samples, the synthesiser, the encoder, the app's own components for the cards and the screens — cost nothing.
| Item | Count | Unit | Total |
|---|---|---|---|
| Storyboard and scenario frames (Gemini) | ≈130 | cents | ≈ €12 |
| Veo 3.1 clips (Google AI Studio) | 10 | €2.85 | €28.50 |
| ElevenLabs Starter, then Creator (Kling 3.0 Pro clips, the music takes) | 2 months | $6 + $22 | $28 |
| Kling 3.0 Pro by API (fal.ai) | 11 | $0.67 | $7.37 |
| Piano samples, synthesiser, ffmpeg, the app's components | free | €0 | |
| Charged for the film | ≈ €75 | ||
| Of which never made the film | 21 of 53 clips | ≈ €30 |