A voice that is actually yours
One 15–30 minute recording session, once. Agent Avatar chunks it, transcribes it, and fine-tunes F5-TTS into a voice you own outright and can use for the rest of your life.
Local voice and video cloning
Stop filming. Start typing.
Agent Avatar clones your voice and your face on your own GPU — then turns a script into video of you delivering it, on demand. No studio. No lighting. No retakes. No uploading your identity to a company that will still have it in ten years.
Your footage, your model, your clips: all of it stays on your drive. Every time.
Coming soon — for NVIDIA GPUs on Linux and WSL2.
# one line, in your own voice and face ./avatar generate \ --voice my-voice \ --face my-face \ --text "Here's what shipped this week." \ -o out/update.mp4 # or the whole script, in one pass curl localhost:8765/lipsync/generate/batch \ -d @release-notes.json # → clips + per-word timestamps, on your drive
Why local
Every cloud avatar service asks you to hand over the two things you can never re-issue: your face and your voice. Then they keep them — on their servers, under their terms, for as long as they feel like it. Agent Avatar never sends them anywhere. Your voice model is a file, in a folder, on your drive. Nobody can retrain on it, leak it in a breach, lose it in an acquisition, or switch it off because your card expired.
One recording session. Unlimited video, in your own voice, for as long as you own the hardware.
Your likeness is getting a fingerprint.
We’re building identity protection straight into the model right now: an invisible signature baked into every avatar you train, and a signed manifest on every clip you render. A video of you will be able to prove it came from you — and a clone that isn’t yours will be tellable from one that is. It will protect your data, your face, and your voice at the file level, not the terms-of-service level.
Your face, cryptographically yours. Shipping soon.
What it does
Voice cloning, talking-head video, batch rendering, and word-level timing — all of it local, all of it scriptable.
One 15–30 minute recording session, once. Agent Avatar chunks it, transcribes it, and fine-tunes F5-TTS into a voice you own outright and can use for the rest of your life.
MuseTalk drives real footage of you. Only have a headshot? SadTalker animates the still.
Batch mode loads the model once and renders every scene in your script in a single pass. Feed it a 40-line VO script and walk away.
Every clip comes back with per-word timestamps. Captions and cuts line up on the first try.
Need it to land at 12.0 seconds? Ask. It pads to length and never clips a word to get there.
Optional face restoration with a blend dial, so you choose how much polish before it tips into uncanny.
Name your voices, faces and source clips, set defaults, and pick a look instead of hunting for file paths.
Every capability is a local HTTP endpoint, and it ships as the first plugin in the Agent Software Suite's plugin API. Wire it into your own pipeline.
How it works
A 15–30 minute audio session, plus a short piece of footage or a single headshot. That is the whole capture step.
Agent Avatar chunks and transcribes the audio and fine-tunes a voice model locally. The result is a file on your drive.
Hand it a line or a forty-scene voiceover script, in the local API or the unified CLI.
Get video of you delivering it, with per-word timestamps for captions — as many times as you like, for as long as you own the hardware.
Interface layer
Everything above is reachable over a local HTTP API, the unified ./avatar CLI, and the Agent Software Suite plugin interface.
Comparison
| Capability | Agent Avatar | HeyGen | Synthesia | D-ID |
|---|---|---|---|---|
| Your face and voice stay on your machine | Yes | No | No | No |
| No per-minute inference fees | Yes | No | No | No |
| Works fully offline | Yes | No | No | No |
| You keep the trained voice model | Yes | No | No | No |
| Runs on your own GPU | Yes | No | No | No |
| Scriptable local API | Yes | Yes | Yes | Yes |
| Batch a whole script in one pass | Yes | Yes | Yes | No |
| Per-word timestamps on every clip | Yes | No | No | No |
Consent policy
Agent Avatar is built to give you control of your own likeness, and that principle does not stop at your own face. You may train a voice or a face model on yourself, or on a person who has given you their explicit, informed permission to do so — and nobody else.
These rules are part of the terms of use. Because generation runs entirely on your machine, enforcement rests with you — which is exactly why we state the line plainly.
FAQ
An NVIDIA GPU with 8GB or more of VRAM, 32GB of system RAM, and roughly 100GB of free disk for the models. Agent Avatar targets Linux and WSL2 today — not native Windows or macOS.
Three conda environments and a multi-gigabyte model download. It is not a one-click installer, and we would rather tell you that up front than have you find out at step four.
Your own, or someone who has given you their explicit consent. Cloning a person without their permission, or using Agent Avatar to impersonate anyone, is prohibited by the terms of use. See the consent policy for the full rules.
On your drive, in folders you choose. Your source footage, your trained voice model, and every rendered clip stay on the machine that made them. Nothing is uploaded for processing.
Only to download the models the first time. After that, generation runs entirely offline.
Agent Avatar is in development and not yet available for purchase or download. Join the Discord to hear when it ships.
Agent Avatar is a standalone product with its own local API and CLI. Agent Studio can drive it as a render backend, so a Studio project can produce narrated video without leaving your machine — but neither one requires the other.
Affiliate Program
Refer paying builders and earn 30% on every invoice they pay for the next 12 months.
30%
Commission
12 invoices
Window
60 days
Cookie
Agent Avatar is coming soon for NVIDIA GPUs on Linux and WSL2. Check what your machine needs, then be first to know when it ships.