Current Model

LongCat Avatar 1.5
LongCat Avatar 1.5

Video Utilities

Audio-driven avatar video from a single photo with sharper lip sync, natural head and body motion, and strong identity preservation (up to 64s).

Quick Navigation

Portrait image *

Source portrait photo of the person to animate (clear face, front-facing works best)

Audio *

Voice or singing track that drives lip sync and performance (trimmed to 64s max)

Prompt (optional)

Field Description: Guide expression, pose, style, or motion (e.g. natural speaking with subtle head movement) Voice tips: To use the voice feature, you need to click on the microphone icon. You need to have a microphone connected to your computer. The voice feature uses 0.5 tokens per audio usage. Current language: English. Please change the language at the bottom if you are speaking in a different language. Multilingual support: GenVR supports multiple languages. If you are prompting in any language besides english, make sure the language setting in the bottom of the page is set to correct language.

Resolution720p

Output resolution: 480p or 720p

Seed-1

Random seed for reproducibility (-1 for random)

Please provide all required parameters (* marked)

20 credits per 5 seconds at 480p, 40 credits per 5 seconds at 720p (min 5s, max 64s)

No Output Yet