Train a custom voice
Prepare a voice corpus, transcribe it, and fine-tune F5-TTS through the guided desktop workflow and CLI.
The product
Train and run the local Avatar workflow directly, then connect it to Agent Studio when the result needs a complete creative-production pipeline.
Inside Avatar
Each capability maps to a real desktop, CLI, or local API workflow.
Prepare a voice corpus, transcribe it, and fine-tune F5-TTS through the guided desktop workflow and CLI.
Use MuseTalk to drive a source video with generated speech for repeatable talking-head scenes.
Use the photo-based workflow when a headshot is the right starting point instead of a source video.
Batch a sequence of scenes and keep word-level timing available for captions and downstream editing.
Apply optional face enhancement and choose how much polish fits the source and the story.
Drive supported workflows from the unified avatar CLI or authenticated local HTTP endpoints.
Name and reuse voice, face, and source entries instead of managing fragile file paths by hand.
When you own both products, Agent Studio can use Avatar capabilities inside a broader editing and production workflow.
Project boundary
Training and generation are designed to run on your machine. Source media and model files are not uploaded for processing.
Account entitlement, updates, and optional connected services may still use network services. That distinction is part of the product, not fine print.
Agent Avatar is in a qualified technical beta. The exact platform matrix and onboarding path are intentionally narrower than a consumer cloud avatar service.
Agent Avatar + Agent Studio
Avatar stays useful as a standalone training and generation app. With Agent Studio, supported voice, face, and lip-sync features can become part of timeline editing, motion, music, review, and final export workflows.
Explore Agent Studio