Most people who make digital work carry a wall down the middle of their head. Code on one side, picture on the other, sound somewhere in a different room, motion bolted on at the end by whoever owns After Effects. Four crafts, four apps, four people, four handoffs. I’ve worked that way for most of my life, and I’ll say the quiet part out loud: the wall was never real. It was a limit of the tools, and the tools are finally catching up.
Here’s the line I keep at the center of everything: code, image, sound, motion — every layer understood, every layer controlled. Read it fast and it sounds like a brag. It isn’t. It’s a description of a single object. A piece isn’t a stack of four deliverables glued together at the deadline. It’s one system with four faces, and the only question that matters is whether one head can hold all of them at once.
01Four crafts, or one object?
The discipline split is an org-chart artifact. Studios separated code, design, audio and motion because no single person could be excellent at all four and no toolchain let you move between them without exporting, re-importing, and losing half the intent in the gap. So the work got cut into lanes, and the seams between the lanes became where the work goes flat.
You can hear those seams. A film where the picture is gorgeous and the sound was clearly chosen from a library afterward. A site where the motion was “added” once the layout was locked, so nothing actually moves with the content. The flatness isn’t a skill problem. It’s a handoff problem. Every wall between disciplines is a place where intent leaks out.
Mood, dynamics, sound — all defined before the first AI generates a single frame.— the rule I build under
When one person holds the whole thing, those decisions stop being sequential. The look isn’t designed and then scored. The cut isn’t timed and then sound-designed. They’re decided together, because in the head that holds all four they were never separate to begin with.
02Sound is built in parallel, never bolted on
This is where I get stubborn, so let me be specific. I build the sound space at the same time as the image — not after. The low end and the cut are designed as one move. A transition lands because the visual change and the audio change were planned as a single event, not because someone found a whoosh that roughly fits.
Sound bolted on at the end always sounds bolted on. You can’t retrofit the relationship between a beat and a frame; you either composed them together or you’re hoping they’ll forgive each other in the mix. Treating audio as the last 5% is the single clearest tell that four people made a thing instead of one system.
- Codestructure & logic
- Imagelook & light
- Soundbuilt in parallel
- Motiontiming & cut
03The console is finally one console
For most of my career, holding all four in one head was a liability — the tools punished you for it. You still had to leave the room every time you crossed a discipline, and the friction of switching cost you more than the breadth bought you. That’s the part that changed. The new generation of AI tools doesn’t make me a better designer or a better sound person. It collapses the distance between the four chairs so the one head that wants to sit in all of them finally can.
The model widens the search space — variants of look, of motion, of tone, faster than any team could. I stay the taste layer and the integration layer: I decide what stays, what gets pushed, and crucially how the four faces line up. The agents are throughput; the coherence is mine. That’s the whole trade. Not “AI makes the work,” but “AI removes the walls that forced the work into four pieces.”
So I don’t think the next generation of creative tools is about speed, and it’s definitely not about replacing the person. It’s about this: one person, directing code, image, sound and motion as a single coherent system, from one console, without the seams. The wall down the middle of the head was always optional. I just finally have the tools to act like it.
