Train a custom voice
Prepare a voice corpus, transcribe it, and fine-tune F5-TTS through the guided desktop workflow and CLI.
Local voice and face production
Your voice. Your face. Your production stack.
Turn a script into a repeatable voice-and-face production workflow. Train and generate locally, then connect Agent Avatar to Agent Studio when the finished video needs a full creative pipeline.
Qualified technical beta — Linux or WSL2 with an NVIDIA GPU. Request access for the current build.
./avatar generate \ --voice my-voice \ --face my-face \ --text "Here's what shipped this week." \ -o out/update.mp4 # batch a script through the local API curl localhost:8765/lipsync/generate/batch \ -d @release-notes.json
See it in motion
Start with the Agent Software Suite overview, then follow along as we publish videos made with Agent Avatar about Agent Avatar.
A short introduction to the wider suite. Dedicated Agent Avatar walkthroughs will be added as they are produced.
Output first
Agent Avatar is a standalone production engine for creators and technical teams. Use it directly, or connect it to Agent Studio when the result needs editing, motion, music, review, and final export.
Prepare a voice corpus, transcribe it, and fine-tune F5-TTS through the guided desktop workflow and CLI.
Use MuseTalk to drive a source video with generated speech for repeatable talking-head scenes.
Use the photo-based workflow when a headshot is the right starting point instead of a source video.
Batch a sequence of scenes and keep word-level timing available for captions and downstream editing.
Apply optional face enhancement and choose how much polish fits the source and the story.
Drive supported workflows from the unified avatar CLI or authenticated local HTTP endpoints.
Name and reuse voice, face, and source entries instead of managing fragile file paths by hand.
When you own both products, Agent Studio can use Avatar capabilities inside a broader editing and production workflow.
How it works
Record the voice, footage, or photo you are authorized to use and prepare the source material for training or generation.
Run the voice and face preparation pipeline on a supported NVIDIA GPU, with progress, checkpoints, and diagnostics.
Submit a line or a batch script through the desktop app, CLI, or local API and produce synchronized output.
Connect Agent Avatar to Agent Studio when the result needs timeline editing, motion, music, review, and final export.
Interface layer
Everything above is reachable over a local HTTP API, the unified ./avatar CLI, and the Agent Software Suite plugin interface.
Local by design
Voice and face preparation runs on a supported local GPU workflow.
Source media and model files are not uploaded for processing; entitlement, updates, and optional integrations may use network services.
Train and generate only with your own likeness or explicit, informed permission from the subject.
The boundary
Train and run voice, face, lip-sync, photo animation, enhancement, catalog, CLI, and local API workflows.
When you have both products, Studio can unlock supported Avatar-powered scenes inside timeline editing, motion, music, review, and export workflows.
Explore Agent StudioComparison
| Capability | Agent Avatar | HeyGen | Synthesia | D-ID |
|---|---|---|---|---|
| Local processing path | Yes | No | No | No |
| Runs on your own GPU | Yes | No | No | No |
| Offline generation after setup | Yes | No | No | No |
| Scriptable local API | Yes | Yes | Yes | Yes |
| Batch a whole script | Yes | Yes | Yes | No |
| Agent Studio connector | Yes | No | No | No |
Consent policy
You may train a voice or face model on yourself, or on a person who has given explicit, informed permission. Do not impersonate people, mislead an audience about synthetic media, or use someone's likeness to obtain money, credentials, or trust.
Read the full consent policyFAQ
It is for creators, educators, developers, and technical teams who need repeatable voice-and-face video and can run a local GPU workflow. Nontechnical self-serve onboarding is not the current target.
The current beta targets Linux or WSL2 with an NVIDIA GPU offering at least 8GB of VRAM, about 32GB of system RAM, and roughly 100GB of storage for models and media.
Training and generation are designed to run on your machine. Source media and model files are not uploaded for processing. Account entitlement, updates, and optional connected services may still use network services.
Yes. Agent Avatar is standalone, and users who have both products can unlock supported Avatar-powered voice, face, and lip-sync workflows inside Agent Studio.
Only yourself or a person who has given explicit, informed consent. Impersonation and deceptive synthetic media are prohibited. Read the consent policy before training or generating.
Not yet. The product is in a qualified technical beta while packaging, onboarding, and production commerce are finalized. Request access for the current availability and support path.
Join the technical beta, or talk to us about a finished Launch Reel for your product.