Skip to main content

Local voice and video cloning

Your face. Your voice. Your machine.

Stop filming. Start typing.

Agent Avatar clones your voice and your face on your own GPU — then turns a script into video of you delivering it, on demand. No studio. No lighting. No retakes. No uploading your identity to a company that will still have it in ten years.

Your footage, your model, your clips: all of it stays on your drive. Every time.

Coming soon — for NVIDIA GPUs on Linux and WSL2.

./avatar generatelocal
# one line, in your own voice and face
./avatar generate \
  --voice my-voice \
  --face my-face \
  --text "Here's what shipped this week." \
  -o out/update.mp4

# or the whole script, in one pass
curl localhost:8765/lipsync/generate/batch \
  -d @release-notes.json
# → clips + per-word timestamps, on your drive
0
clips uploaded anywhere
renders per month
1
recording session, once
3
engines: voice, video, photo

Why local

Nobody can steal a clone that never leaves your desk.

Every cloud avatar service asks you to hand over the two things you can never re-issue: your face and your voice. Then they keep them — on their servers, under their terms, for as long as they feel like it. Agent Avatar never sends them anywhere. Your voice model is a file, in a folder, on your drive. Nobody can retrain on it, leak it in a breach, lose it in an acquisition, or switch it off because your card expired.

One recording session. Unlimited video, in your own voice, for as long as you own the hardware.

Identity Shield

In development now

Your likeness is getting a fingerprint.

We’re building identity protection straight into the model right now: an invisible signature baked into every avatar you train, and a signed manifest on every clip you render. A video of you will be able to prove it came from you — and a clone that isn’t yours will be tellable from one that is. It will protect your data, your face, and your voice at the file level, not the terms-of-service level.

Your face, cryptographically yours. Shipping soon.

What it does

A studio that fits on the GPU you already own.

Voice cloning, talking-head video, batch rendering, and word-level timing — all of it local, all of it scriptable.

V

A voice that is actually yours

One 15–30 minute recording session, once. Agent Avatar chunks it, transcribes it, and fine-tunes F5-TTS into a voice you own outright and can use for the rest of your life.

F

Video from footage — or a single photo

MuseTalk drives real footage of you. Only have a headshot? SadTalker animates the still.

B

Built for whole scripts, not party tricks

Batch mode loads the model once and renders every scene in your script in a single pass. Feed it a 40-line VO script and walk away.

T

Word-perfect timing, automatically

Every clip comes back with per-word timestamps. Captions and cuts line up on the first try.

R

Hit an exact runtime

Need it to land at 12.0 seconds? Ask. It pads to length and never clips a word to get there.

E

Looks human, not plastic

Optional face restoration with a blend dial, so you choose how much polish before it tips into uncanny.

L

A library of looks

Name your voices, faces and source clips, set defaults, and pick a look instead of hunting for file paths.

A

Fully scriptable

Every capability is a local HTTP endpoint, and it ships as the first plugin in the Agent Software Suite's plugin API. Wire it into your own pipeline.

How it works

Record once. Render for as long as you own the machine.

1

Record yourself once

A 15–30 minute audio session, plus a short piece of footage or a single headshot. That is the whole capture step.

2

Train on your own GPU

Agent Avatar chunks and transcribes the audio and fine-tunes a voice model locally. The result is a file on your drive.

3

Write the script

Hand it a line or a forty-scene voiceover script, in the local API or the unified CLI.

4

Render and reuse

Get video of you delivering it, with per-word timestamps for captions — as many times as you like, for as long as you own the hardware.

Voice cloning (F5-TTS)
Lip sync (MuseTalk)
Photo animation (SadTalker)
Face restoration
Voice and face catalog
Local HTTP API + CLI

Interface layer

Everything above is reachable over a local HTTP API, the unified ./avatar CLI, and the Agent Software Suite plugin interface.

Comparison

The difference is where your face lives.

Your face and voice stay on your machine

Agent AvatarYesHeyGenNoSynthesiaNoD-IDNo

No per-minute inference fees

Agent AvatarYesHeyGenNoSynthesiaNoD-IDNo

Works fully offline

Agent AvatarYesHeyGenNoSynthesiaNoD-IDNo

You keep the trained voice model

Agent AvatarYesHeyGenNoSynthesiaNoD-IDNo

Runs on your own GPU

Agent AvatarYesHeyGenNoSynthesiaNoD-IDNo

Scriptable local API

Agent AvatarYesHeyGenYesSynthesiaYesD-IDYes

Batch a whole script in one pass

Agent AvatarYesHeyGenYesSynthesiaYesD-IDNo

Per-word timestamps on every clip

Agent AvatarYesHeyGenNoSynthesiaNoD-IDNo

FAQ

Common questions

What hardware do I need?

An NVIDIA GPU with 8GB or more of VRAM, 32GB of system RAM, and roughly 100GB of free disk for the models. Agent Avatar targets Linux and WSL2 today — not native Windows or macOS.

What does the install actually involve?

Three conda environments and a multi-gigabyte model download. It is not a one-click installer, and we would rather tell you that up front than have you find out at step four.

Whose face and voice am I allowed to clone?

Your own, or someone who has given you their explicit consent. Cloning a person without their permission, or using Agent Avatar to impersonate anyone, is prohibited by the terms of use. See the consent policy for the full rules.

Where are my recordings and models stored?

On your drive, in folders you choose. Your source footage, your trained voice model, and every rendered clip stay on the machine that made them. Nothing is uploaded for processing.

Does it need an internet connection?

Only to download the models the first time. After that, generation runs entirely offline.

When can I get it?

Agent Avatar is in development and not yet available for purchase or download. Join the Discord to hear when it ships.

How does it relate to Agent Studio?

Agent Avatar is a standalone product with its own local API and CLI. Agent Studio can drive it as a render backend, so a Studio project can produce narrated video without leaving your machine — but neither one requires the other.

Affiliate Program

Promote Agent Software Suite. Earn 30% recurring.

Refer paying builders and earn 30% on every invoice they pay for the next 12 months.

30%

Commission

12 invoices

Window

60 days

Cookie

One recording session away from never filming again.

Agent Avatar is coming soon for NVIDIA GPUs on Linux and WSL2. Check what your machine needs, then be first to know when it ships.