Generate or source voice-over, music, sound effects, avatars, captions, stock video, and stock audio, place each result at the correct scene or master time, then mix, trim, and verify synchronization on the timeline.
Where: AI assistant (video project) → describe the audio/captions — Ask the AI for VO / music / SFX / captions
Key ideas
- Scene and master timing differ: Scene voice-over can be generated and fitted to one scene; general audio layers can be placed at an explicit master-timeline time.
- Generation and placement are separate: Some tools create persistent media while placement tools add and trim it on the timeline.
- Captions need an audio source: Speech-synchronized captions use the chosen audio layer or mixed speech source rather than guessing from visuals.
Steps
- Select the target scene or note the intended master start time.
- Provide final narration text, voice intent, music mood/duration, or SFX description.
- Choose a voice when needed, then generate voice, music, SFX, avatar, or locate stock media.
- Preview the generated asset before placement and regenerate if content or duration is wrong.
- Add it to the intended scene or master-timeline lane and trim/fade it to fit.
- Generate captions from the correct speech source and inspect line breaks and timing.
- Play across scene boundaries, balance levels, prevent clipping, and verify captions after any audio edit.
Available media actions
- Generate voice and add voice-over to scene
- Generate music or SFX
- Generate avatar
- Generate captions
- Add audio to timeline
- Search Pexels video or stock audio
Tips
- Finalize narration wording before captions.
- Leave music headroom under speech.
Limitations and important notes
- Generated duration and pronunciation can vary.
- Moving or trimming speech after captioning can invalidate caption timing.
Troubleshooting
Captions are out of sync.
Select the final speech layer, regenerate captions after trims, and verify the master start time.