Home
Dubb

Do AI Avatar Videos Actually Convert?

Ruben

Ruben

Avatar videos work when the synthetic face is not the main thing on screen. A full-frame talking avatar invites viewers to look for flaws and usually find them. Placing a small avatar over a screen, a profile, or a property listing shifts attention to the content, which is where these videos perform best today.

What Is an AI Avatar Video?

An AI avatar video is a video in which a synthetic presenter, generated from a photograph or created from a description, delivers a script in a synthesised or cloned voice. It differs from a recorded video in that no camera is involved, so the same presenter can deliver unlimited scripts without anyone sitting down to film.

The workflow described here runs through Dubb Agent connected to HeyGen, which handles the avatar and voice generation. The agent writes the script, produces the video, and uploads the result to a shareable landing page, so the whole sequence happens in one conversation rather than across several tools.

Table of Contents

  1. Why Avatar Videos Usually Fail
  2. The Question Nobody Asks: Compared to What?
  3. The Sandwich Video
  4. Where Avatar Videos Genuinely Work
  5. Avatar Screen Videos: The Format That Works
  6. What You Need to Set It Up
  7. Recording a Voice Sample
  8. Creating or Choosing an Avatar
  9. Expect to Iterate
  10. Consent, Disclosure, and Likeness
  11. What It Costs
  12. Publishing and Repurposing
  13. Avatar Video Tools Compared
  14. Common Mistakes and Troubleshooting
  15. Best Practices I Actually Follow
  16. Proof: Why This Actually Works
  17. Frequently Asked Questions
  18. Pick the Format Before the Tool

Why Avatar Videos Usually Fail

I want to lead with the problem rather than the pitch, because the failure mode here is specific and worth naming.

The roboticist Masahiro Mori described it in 1970 as the uncanny valley: as something becomes more human-like, our affinity for it rises, and then drops sharply in the region just short of convincing, before recovering once the likeness is complete. Avatar videos sit in that dip. Believability climbs while you watch, and then something in the mouth or the blink or the evenness of the delivery gives way, and the feeling that replaces it is not neutral. It is mild revulsion.

That matters commercially because the drop happens mid-message. Your prospect is not evaluating your offer at that point, they are working out what is wrong with your face. Whatever came next is lost.

The specific tell is a combination of too perfect and not perfect enough. The skin is too even, the delivery too metronomic, and then the timing of a blink is slightly wrong. None of it is individually obvious. Together it reads as off.

So the request we hear constantly, which is roughly "I am busy and not great on camera, let the AI make personalised videos for my prospects," is a reasonable thing to want and usually not the thing that will get results. The honest answer is that the technology is not there yet for a full-frame synthetic face addressing someone by name.

The Question Nobody Asks: Compared to What?

Here is the comparison that decides it, and almost nobody makes it explicitly.

The choice is rarely between a personalised avatar video and nothing. It is between a personalised avatar video and a properly produced evergreen video that you make once and reuse: edited, with b-roll, music, testimonials, and social proof in it.

Put those side by side from the recipient's point of view. One is addressed to them by name and looks slightly wrong. The other is generic but well made, carries actual evidence, and does not trigger anything unsettling.

Speaking personally, I would rather receive the produced generic video. The personalisation buys less than people think, and the discomfort costs more than they expect. Personalisation is only an advantage when everything else holds up.

Which reframes the decision. The question is not "should I use avatar videos," it is "what is the best thing I can send this person for the effort I have available," and a good evergreen video usually beats a mediocre personalised one.

The Sandwich Video

There is a third option that beats both, and it is the one I would try before any of this.

Record a real personalised introduction, a genuinely short one. Six seconds is enough: their name, one specific reason you are reaching out, and a handoff into the produced piece. Something close to "Hi Darius, made this for you, would love your feedback." Then let it play into the evergreen video you already made.

That is the sandwich. Real human on the front, produced content in the middle, and the personalisation is authentic because it is you.

Six seconds is the part people miss. The reason personalised video does not scale is that people try to personalise the whole message. You do not need to. You need the first few seconds to be unmistakably for this person, and everything after that can be the same for everyone.

If you can find ten minutes to batch-record thirty of those, you have thirty personalised sends and no synthetic likeness in the chain at all.

Where Avatar Videos Genuinely Work

None of which means avatar videos have no place. There are applications where they are genuinely good, and they share a property: nobody is expecting intimacy.

Training material. Internal onboarding, process walkthroughs, compliance modules. The viewer wants the information and is not looking for a relationship.

Support and help content. Explaining a feature, answering a common question. Consistent delivery is an advantage and updating one script is easier than refilming.

Generic or high-volume content. Anything you would otherwise not make at all because filming it is not worth the time.

Multilingual versions. The same content across languages, where the alternative is subtitles or nothing.

The test I apply is whether I would be embarrassed if the recipient realised it was an avatar. For a training module, not at all. For a first approach to a prospect I want a relationship with, considerably.

Avatar Screen Videos: The Format That Works

This is the part worth the article, and it came out of trying the obvious thing and not liking it.

Having generated a full-frame talking avatar and found it flat, the move that fixed it was to stop making the face the subject. Put a screen behind the avatar, shrink the avatar into a corner, and let the content carry the video: a prospect's website, their profile, a property listing, a document, a slide.

We ended up calling these avatar screen videos, and the reasoning is simple. When a face fills the frame, the eyes go looking for imperfections and find them. When the face occupies a small part of the frame over something worth looking at, attention goes to the content, the delivery, and the story. The avatar becomes incidental rather than the thing under inspection.

Think of it as the same reason a small stain ruins an expensive suit. Attention goes to the flaw, not the quality around it. Give the eye something better to look at and the flaw stops being the subject.

The practical upside is that this costs no extra effort. Generating an avatar screen video takes about as long as generating a full-frame one. You can even ask the agent to capture a screenshot of a prospect's site or profile to use as the background, which it does by driving a browser in the background.

So the choice between the two formats is genuinely free, and one of them works better. That is an unusually easy decision.

What You Need to Set It Up

Three things, and the first two are quick.

A Dubb Agent account. There is a free trial without a credit card. Generation consumes AI credits, because producing video this way is computationally expensive.

A HeyGen account, connected. Go to integrations in Dubb Agent, find HeyGen, and connect. It needs an API key from your HeyGen account. If you get stuck, ask the agent directly how to connect HeyGen and it will walk you through it; keep the integrations page open in another tab so you can paste the key across.

The avatar video generator skill, turned on. In the skills section, switch it on. This is what teaches the agent the whole sequence, from generating the avatar through to producing and uploading the finished video. Without it you will get much worse results.

Then it is a conversation. Upload a photo, upload a voice sample, and describe what you want.

Recording a Voice Sample

The voice sample matters more than the photo, because a monotone sample produces a monotone video and no amount of scripting rescues it.

You do not need equipment. A modern phone's voice memo app, held about six inches from your mouth, in a quiet room, is enough. A USB microphone is better if you have one, but the difference is smaller than the difference between a quiet room and a noisy one.

Factor Do This Why It Shows Up Later
Room Quiet and soft. A car or a clothes-filled closet works as a makeshift booth Background noise and room echo get cloned along with your voice
Distance About six inches from your mouth Too far picks up the room, too close distorts
Range Vary it deliberately. Go quiet, then emphatic, within the same sample The clone can only reproduce dynamics it heard. Flat in, flat out
Emotion Speak with actual interest in what you are saying Reading a sample dutifully produces a voice that sounds dutiful forever
Length A few sentences, not a single line More material gives the clone more to work from

If cloning your own voice feels like a step too far, choose one from the stock library instead. We used a library voice for one of our best results, picked for warmth rather than for sounding like anyone in particular.

Creating or Choosing an Avatar

Three routes, in increasing order of interest.

Upload your own photo. The agent builds an avatar from it. Simplest, and it is you.

Pick from the stock library. Fast, and nobody has to consent to anything. The drawback is that stock avatars appear in other people's videos too.

Describe one and have it generated. This is the one I would point people at. Someone on our team built ours by describing the features they wanted, and the result was a presenter designed to our specification rather than chosen from a shelf.

A custom-generated avatar avoids both the consent question and the stock-library problem. Worth reading the platform's terms on what rights you have over a generated likeness and how it may be used, since that varies and it is the sort of thing you want to know before you build a channel around one.

Expect to Iterate

The first output will not be the one you use, and that is normal rather than a failure.

Ours came back flat and, oddly, too cheerful. The fix was to tell the agent so, in plain language: stop smiling so much, it reads as awkward. Version two was noticeably better. Later rounds fixed a background that showed a URL rather than the site itself, which took a couple of goes to explain.

The useful habit is to describe the problem rather than prescribe the fix. "Too smiley and it feels uncanny" gets you further than trying to specify facial parameters, because the agent is better at interpreting a critique than you are at writing a specification for a face.

Budget three or four rounds. It is still faster than filming.

This section is not in the source training, and it belongs before you scale any of this rather than after.

Whose face and voice. Cloning yourself is uncomplicated. Cloning a colleague needs their explicit agreement, ideally written, and it should cover what the avatar may be used for and what happens when they leave. Cloning anyone who has not agreed is a different category of problem: most jurisdictions protect a person's likeness and voice from commercial use without permission, and several have recently strengthened those protections specifically for AI-generated replicas. Do not build a prospecting campaign on someone else's face.

Whether to say it is AI. The use case that makes people uneasy is a synthetic version of a real person saying a prospect's name, sent to that prospect, with nothing indicating it is generated. Whatever the rules say where you operate, consider how the recipient will feel when they work it out, because they often will. Discovering that a warm personal message was synthetic tends to cost more trust than the message ever earned.

Disclosure requirements are also moving quickly. Some jurisdictions now require AI-generated media to be labelled in certain contexts, and the picture differs by country and by state. Check what applies where you and your recipients are.

The version I would default to. Use a custom generated avatar that is not a real person, or be open about it. A branded synthetic presenter that nobody mistakes for a colleague raises none of these questions, which is a large part of why it is my preferred route.

What It Costs

Generation consumes credits, and the cost varies with video length, how many revisions you run, and the quality of voice you choose.

As a rough guide, expect somewhere around one to three dollars per video. Confirm that against current credit pricing before planning a large run, because it is the kind of figure that moves.

The useful way to read it is per outcome rather than per video. At a couple of dollars, a video that saves a meeting is trivially worth it and a thousand videos into a cold list is a real budget. And the alternative is not free either. If a couple of dollars a video is more than you want to spend, the answer is batch recording: block an hour, record thirty short personalised openers, and you have covered the same ground with your actual face.

Publishing and Repurposing

Once a video exists, ask the agent to upload it to your Dubb account, which generates a landing page you can share anywhere.

From there, any video on the platform can be pushed to a connected YouTube account, as a short or as a long-form upload. That makes powering a channel a matter of generating and publishing rather than filming and editing.

Beyond YouTube, the usual repurposing route applies: tools that take a YouTube upload and schedule it out to Instagram, TikTok, Facebook and X. We covered that workflow in the post on repurposing social media content across every channel.

A caution that follows from the first half of this article: scaling distribution multiplies whatever you made. If the video sits in the uncanny valley, publishing it to five channels does not fix it, it just shows more people. Get the format right before you turn on the volume.

Avatar Video Tools Compared

The generation layer is a competitive category. What differs is what happens either side of the generation.

Tool What It Handles What Happens After Best Fit
Dubb Agent with HeyGen Script, avatar, voice, screen background, and the video, in one conversation Lands on a hosted, trackable, sendable page; publishes to YouTube Sales teams who need the video sent and measured, not just made
HeyGen on its own Avatars, voice cloning, and translation across many languages You export and distribute it yourself People who want the generation layer directly
Synthesia Avatar video built around training and internal communication Shared inside its own workspace or exported Enterprise training and multilingual internal content
D-ID Digital avatars plus interactive conversational agents Embedded in your own product or site Teams building an avatar into an interactive experience

If all you need is the file, go to a generation tool directly. The argument for running it through an agent is that the script, the background, the production and the upload happen in one place, and the finished video arrives somewhere it can be sent and tracked rather than in a downloads folder.

Common Mistakes and Troubleshooting

Making the face the subject. The single biggest one. Put the avatar over a screen and shrink it.

A flat voice sample. Monotone in, monotone out. Vary your delivery deliberately while recording.

Recording the sample in a noisy room. The noise gets cloned too. Use a car or a closet if you have to.

Accepting the first generation. Expect three or four rounds and give feedback in plain language.

Prescribing fixes instead of describing problems. "This feels uncanny and too smiley" works better than trying to specify facial settings.

Using someone else's face or voice. Get explicit permission, in writing, and never for someone who has not agreed.

Using an avatar where a real six seconds would do. The sandwich video beats a synthetic one for first contact almost every time.

Scaling before the format is right. Distribution multiplies the video you made, including its problems.

Best Practices I Actually Follow

Default to avatar screen videos. Same effort, better result, fewer places for the eye to snag.

Use a custom generated avatar rather than a real person. It removes the consent question and the stock-library problem at once.

Try the sandwich first. Six real seconds plus a produced evergreen video is hard to beat, and it costs nothing to generate.

Reserve avatars for content where nobody expects intimacy. Training, support, explainers, translations.

Record one good voice sample and stop. It is reused indefinitely, so ten minutes on getting it right pays out for a long time.

Watch it as a stranger before sending it. If you feel the dip, they will too.

Proof: Why This Actually Works

The mechanism is attentional rather than technical. A face at full frame is inspected, because faces are what people are best at reading. The same face at a tenth of the frame, over content worth looking at, is not inspected at all.

Two patterns hold consistently across the people we help. The first concerns format. Avatar videos where the synthetic presenter is small and the screen carries the content are kept and sent; full-frame avatar videos are generated, watched once by the person who made them, and quietly not sent. The maker's own discomfort is the most reliable predictor there is.

The second concerns the voice sample. People who record a few sentences with real variation in them end up using the resulting voice repeatedly, and people who read a sample flatly abandon the clone and switch to a library voice. The sample is a one-time decision with a long tail.

Methodology note: these are directional observations drawn from aggregated, anonymized usage patterns across Dubb users, not a controlled study. No figures are attached to either pattern, and results vary by audience, use case, and how the videos are distributed.

What I take from it is that the format decision does more work than the tool decision. The same generation engine produces something you will send or something you will not, depending entirely on how much of the frame the face occupies.

Frequently Asked Questions

Do AI avatar videos work for sales outreach?

Usually not for first contact. A full-frame synthetic face addressing someone by name tends to fall into the uncanny valley, and the drop in believability happens mid-message, so whatever you said next is lost.

What does work is either a small avatar placed over a screen so the content carries the video, or a six-second real recording of you introducing a produced evergreen video. For a stranger you want a relationship with, the real six seconds is hard to beat.

What is the uncanny valley?

A term coined by the roboticist Masahiro Mori in 1970 describing how our affinity for something rises as it becomes more human-like, then drops sharply just short of convincing, before recovering once the likeness is complete.

Avatar videos currently sit in that dip. The giveaway is usually a combination of too perfect and not perfect enough: skin that is too even, delivery that is too regular, and then a blink whose timing is slightly wrong.

What is an avatar screen video?

A video where a small synthetic presenter is placed over a screen, such as a prospect's website, a profile, a property listing, or a slide, rather than filling the frame.

It works because attention goes to the content rather than to the face. A face at full frame gets inspected and the flaws get found; the same face occupying a small part of the frame becomes incidental. It takes no longer to produce than a full-frame avatar video, which makes it an easy default.

How do I record a voice sample for cloning?

Use your phone's voice memo app, held about six inches from your mouth, in a quiet room. A car or a closet full of clothes works well as a makeshift booth if your space is noisy.

Record a few sentences rather than one line, and vary your delivery deliberately: go quiet, then emphatic, within the same sample. The clone can only reproduce dynamics it heard, so a flat sample produces a flat voice in every video you ever make with it.

Do I need permission to clone someone's voice or face?

Yes. Cloning yourself is uncomplicated, but cloning a colleague needs their explicit agreement, ideally written, covering what the avatar may be used for and what happens if they leave.

Cloning anyone who has not agreed is a different category of problem, since most jurisdictions protect a person's likeness and voice from commercial use without permission, and several have strengthened those protections specifically for AI replicas. Using a custom generated avatar that is not a real person avoids the question entirely.

How much does an AI avatar video cost?

Generation consumes credits, and the cost depends on video length, how many revisions you run, and the voice quality you choose. As a rough guide, expect around one to three dollars per video, and confirm against current pricing before planning a large run.

Read it per outcome rather than per video. At those numbers, one video that saves a meeting pays for many, while a thousand videos into a cold list is a real budget decision.

Pick the Format Before the Tool

The sequence that works: connect HeyGen to Dubb Agent, turn on the avatar video generator skill, record one good voice sample with real range in it, generate a custom avatar rather than cloning a colleague, and put that avatar small over a screen rather than large in the middle.

Then, before you send anything, watch it as though you received it. If you feel the dip, so will they, and the fix is almost always to make the face smaller rather than to regenerate it again.

Dubb Agent handles the whole chain, from the script through the generation to a landing page and a YouTube upload, which is what makes iterating cheap enough to actually do. The format judgment is still yours, and it matters more than the tool.

If you change one thing after reading this, try the sandwich video before you build an avatar at all. Six real seconds in front of a well-made evergreen video is still the most reliable thing in this article, and you already have everything you need to make it.

About the Author: Ruben
Ruben

CEO and Founder of Dubb and host of Connection Loop (a top 3% worldwide podcast), author of the best-selling book Click Record. Passionate about helping people succeed with video, AI, and automation. Empowers businesses to grow through innovative technology and storytelling.

Related Posts

View all posts