Video artworksGenerative AI artist
How I Built an 8-Part AI Avatar Healthcare Video Series
A behind-the-scenes case study on building eight AI-avatar healthcare videos through art direction, medical visuals, motion design and a hybrid production workflow.
RELEASED
7/20/20263 min read


Eight Videos, One AI Presenter and a Production Workflow Built Through Failure
Eight 90-second healthcare videos may sound like a straightforward AI-avatar project: generate the presenter, add several medical illustrations and edit everything together.
That was the original assumption.
In practice, I had to build an entire production system around the avatar: visual development, medical B-roll, motion design, reference-based generation, stock footage, compositing and quality control.
The result was twelve minutes of branded healthcare content—and a workflow that eventually reduced my production time from approximately eight hours to four hours per later episode.
The first avatar was not ready for release
This was my first substantial client project built around a generated presenter.
We started with Synthesia, but the initial result was a disappointment. The movement felt disconnected, gestures were limited and the presenter did not carry the professional authority expected in healthcare communication.
The problem was bigger than appearance. In medical content, unnatural movement can reduce trust.
We improved the source recording, tested different angles, controlled the wardrobe and framing, explored generated presenter environments and reduced the amount of time the avatar had to remain continuously visible.
The final solution was not to make the avatar do everything. It became one layer inside a larger edit.
Designing a visual language for eight films
Each video needed to explain processes such as blood-cell analysis, kidney filtration, iron storage, immunity, thyroid function and body composition.
The visuals had to remain understandable without becoming either frightening or scientifically misleading.
I developed a shared design system using:
- cool white and grey backgrounds;
- graphite and dark metal;
- turquoise and restrained coral accents;
- translucent medical forms;
- soft glass interfaces;
- Poppins typography;
- calm, non-alarming motion.
This system allowed every episode to have its own visual metaphors while remaining part of the same series.
A beautiful medical error is still an error
Image models can produce convincing diagrams without understanding what they mean.
During production, I encountered incorrect electrolyte arrows, unreliable anatomical structures, unsuitable examination procedures and generated interfaces that looked professional but communicated the wrong relationship.
My role therefore extended beyond prompting. Every visual had to be checked as communication: What does the viewer understand from this image, and is that actually what the script says?
When precise biological direction was too risky, I simplified the visual into balance, filtration, storage or measurement rather than pretending to show a detailed mechanism.
The end frame became the beginning
One of the most useful methods developed during the project was designing the end frame first.
The approved end frame established:
- composition;
- hierarchy;
- colour;
- final medical state;
- object count;
- the exact destination of the animation.
The start frame was then derived from it by removing only the completed action: particles, illumination, a filled storage unit or an activated measurement.
This gave Kling a controlled path between two approved states instead of asking it to invent the scene while animating it.
Choosing the tool according to the problem
The final workflow was deliberately hybrid:
- Synthesia and HeyGen were explored for the presenter;
- ChatGPT and Nano Banana supported visual development and reference-based image editing;
- Kling handled contained start-to-end-frame motion;
- stock footage replaced AI when realism was faster and more reliable;
- Codex and Remotion were introduced for repeatable programmatic animation;
- Final Cut Pro remained the centre of assembly, masking, typography, timing and correction;
- After Effects was reserved for tasks that required more precise motion control.
The important automation was not “generate everything.” It was knowing what each tool should and should not be asked to do.
What Codex changed
I originally started researching Codex because I wanted to understand whether coding could accelerate motion-design work.
Instead, I discovered a different way to collaborate with animation: describe components, states and timing, then inspect and revise the generated Remotion composition.
It was not automatically successful. My first coded blood-cell animation contained sequencing glitches and unstable movement. I eventually rebuilt part of it manually in Final Cut.
But later, once scenes were separated and instructions became modular, Codex made repeated animation structures substantially faster to produce and revise.
In my own production comparison, later episodes dropped from approximately eight hours to four. This is a working estimate rather than a controlled benchmark, but the difference was practical and visible.
Budget lesson
The first package was priced at approximately one dollar per finished second, with generation credits added separately.
That was a market-entry price, not a sustainable production benchmark.
The real cost included:
- avatar preparation;
- visual development;
- image and video generation;
- stock licensing;
- motion design;
- editing;
- medical and visual quality control;
- client revisions.
The project helped me understand that AI can reduce implementation time, but it does not remove art direction, editorial judgement or quality control from the budget.
The result
The final delivery was not simply an avatar speaking over generated B-roll. It became a reusable content system: one presenter, one visual language and eight specialised healthcare explainers.
I now offer this workflow to businesses that need presenter-led educational, product or expert content without organising a new physical shoot for every episode.
Need an AI presenter or a complete branded video series?