1. Upload a face video
Choose a clip with one clear speaker, a visible mouth, steady framing, and enough facial detail for the model to follow.
Add a video and new audio
Choose a clear face video first. The preview will show your source clip, then the completed lip-synced result.
Preview
Source video
Recent history
MoreNo recent edits yet.
AI Lip Sync Online
YuzuTime AI Lip Sync matches a visible speaker's mouth movement to the speech in a new audio file. Upload both files in the workspace, review the generated video, and download the result.

Clear, audio-based pricing
AI Lip Sync is available to logged-in users with a paid account history. The server rounds the audio duration to the nearest second and charges 4 credits per billed second. The workspace shows an estimate before submission and the confirmed charge after the task is created.
Choose a clip with one clear speaker, a visible mouth, steady framing, and enough facial detail for the model to follow.
Upload spoken audio with a clean voice and limited background noise. The audio duration determines the credit cost.
Confirm the estimated cost, create the task, wait for processing, then review and download the completed video.
Lip-sync quality depends on what the model can see and hear. Prepare the source files before generating instead of relying on broad promises such as perfect or flawless results.
Synchronize an existing face video to a separately prepared voice track. This tool changes mouth movement; it does not translate or create the audio.
Update a presenter clip with revised narration when you have permission to use the video and voice.
Test alternate voice tracks for short promotional or social clips while keeping the original scene.
This workflow requires both a video URL and an audio URL after upload. Results can vary with angle, occlusion, compression, fast speech, and the number of visible people. Review every output before publishing.
Strong side profiles or a mouth covered by a hand, microphone, hair, or shadow can create distorted or delayed-looking movement.
The current workflow does not offer speaker selection. Use a clip with one obvious speaker; multi-person scenes may be unreliable.
YuzuTime synchronizes the video to the audio you provide. Translate, record, or produce the replacement voice track before uploading it.
The new audio drives the billed duration. Prepare and trim both files so the useful sections align, then check the entire result.
Only upload content you own or are authorized to use. Obtain consent before changing a recognizable person's apparent speech or using someone else's voice. Do not use generated media for deception, impersonation, harassment, or rights infringement. Uploaded files are processed to provide the requested generation; consult the current policies for handling and account terms.
AI lip sync changes the visible mouth movement in a video so it follows the speech in a separate audio file. You provide the source face video and the replacement audio; the generated result is returned as a video.
Upload a clear face video, upload the new spoken-audio file, review the estimated credit cost, and submit the task. YuzuTime uploads both files, creates an asynchronous generation, and shows the video when processing finishes.
No free or no-sign-up promise is made for this tool. You must log in, have a paid account history, and have enough credits. The server charges 4 credits for each rounded second of uploaded audio.
The server reads the audio duration, rounds it to the nearest whole second, and multiplies that number by 4. The browser estimate is informative; the server-confirmed duration and credits are authoritative.
Use a browser-readable video with one clear, visible face and a separate spoken-audio file with little noise or overlap. Strong side angles, covered mouths, low resolution, multiple people, and noisy audio can reduce quality.
No. This tool synchronizes mouth movement to the audio you upload. It does not translate the original speech or generate a replacement voice track.
The current workspace does not provide speaker selection or speaker-to-audio mapping. For more predictable results, use a clip with one obvious speaker and keep other visible faces out of frame.
No. The current AI Lip Sync endpoint requires a video and an audio file. Use the Image to Video tool first if you need to create motion from a still image, then review whether the resulting clip is suitable.
That depends on your rights to the source video, face, voice, audio, and any other protected material, as well as the applicable YuzuTime terms. Do not treat the tool as permission to use someone else's identity or content.
Common causes include background noise, fast or overlapping speech, low-resolution faces, strong side angles, mouth occlusion, compression, or multiple visible people. Try a clearer front-facing clip and cleaner, trimmed speech audio.
Files are uploaded and processed to create the requested lip-sync task. Avoid confidential or sensitive content and read the current Privacy Policy for YuzuTime's information-handling details.
Upload a clear face video and spoken audio, review the estimated cost, and generate a new synchronized video.
Open AI Lip Sync