Provide Two Corresponding Recordings
Upload the video and the separate audio captured for the same event or performance.
Align a separate audio recording with the matching video footage. The Pexo AI video agent uses matching content or a supplied cue to set the initial offset and presents the aligned result for drift review.
Matching recordings
Add one video and the audio recording that belongs with the same event.
A clap, spoken phrase, or other supplied cue can help identify the intended sync point.
What it is
Sync Audio and Video corrects the starting timing difference between a video recording and a separate audio recording of the same event.
The workflow aligns existing recordings; it does not generate new speech or change visible mouth movement.
Upload the video and the separate audio captured for the same event or performance.
Use matching speech, a clap, or another shared moment to establish the intended starting point.
Check the start and later moments to see whether picture and sound stay together.
How it works
Provide the corresponding recordings, establish their shared cue, and review both initial offset and later drift.
Add one video and the separate audio track that corresponds to the same action or speech.
The Agent compares matching content or follows the cue you provide to position the audio start.
Review the beginning and later sections, then request an offset adjustment if the sound moves away from the picture.
Agent decisions
The Agent connects the matching sources while leaving final timing judgment to playback review.
Uses matching content or a supplied cue to identify the intended shared moment.
Positions the separate audio relative to the video timeline.
Follows your start and end instructions when the two recordings have different lengths.
Presents the aligned result so you can inspect timing at several points.
Common questions
Provide one video and one separate audio recording that capture the same event, speech, or performance.
The Agent can use matching content or a cue you identify, such as a clap or spoken phrase.
Initial offset is the timing difference at the start; drift is a timing difference that grows later in the recording.
Check a visible or audible cue near the beginning, then inspect later moments to see whether timing still matches.
No. This workflow aligns recordings that already correspond; lip sync changes visible mouth movement to follow audio.
Start with matching sources
Upload both recordings, identify a shared cue, and review timing across the result.