A silent reference can show mouth movement, gestures and visual rhythm, but it does not establish the original words, music or sound effects. Describe those visible actions and mark audio as unknown. If you add dialogue or a soundtrack concept, label it as a new creative proposal rather than a recovered transcript. Do not promise exact speech from visual analysis alone.
Separate three kinds of information
First, record what the image sequence visibly supports: a person opens their mouth, pauses, points and smiles. Second, note any independently available audio that was actually heard or transcribed through a suitable authorized workflow. Third, list proposed sound or dialogue for a new version. These categories should not collapse into one confident narrative.
A reaction can suggest surprise without proving someone said a particular sentence. A visible drum performance does not identify the recorded music track. A rhythmic edit may inspire musical timing, but a visual pulse is not a measurement of a soundtrack's tempo unless the soundtrack itself is available and checked.
Write a visual-only brief honestly
Analyze this prepared video as visual evidence only. Describe the visible subject actions and their order. If a person appears to speak, record visible mouth movement without inventing words. Mark original dialogue, music and sound effects as unknown unless separately supplied and verified. Put any proposed sound design in a clearly labeled creative section, not in the observed reconstruction.
This example is useful even when a service accepts a video file, because acceptance of a container does not tell you which modalities the analysis pipeline uses. Check the controls and returned evidence. Do not assume an audio transcript exists merely because the source file originally contained audio.
Add a new soundtrack as a separate production decision
For a new short film, you can propose footsteps, a door chime or a short musical accent without claiming those sounds were present in the reference. Make the proposal specific: what event motivates the sound, where it begins and whether it continues across a cut. Avoid creating unexplained speech for a real person in a factual-looking context.
If you want spoken dialogue, write and approve it before attempting any supported audio or video workflow. Confirm the voice rights and the intended portrayal. External narration or editing may be needed; Imagild's visual reverse-prompt function should not be advertised as an automatic voice-cloning, transcription or lip-reading service.
- Confirm whether the reviewed material includes usable audio evidence.
- Describe visible speech-like motion without transcribing it.
- Keep observed events and proposed dialogue in separate sections.
- Give each proposed sound a visible cause or editorial purpose.
- Check the actual generated clip for audio rather than assuming it.
- Review words and portrayal before publishing a voiced version.
Keep prompt roles clear
Runway's image-to-video prompting guide treats action and temporal progression as core visual instructions. That is helpful context for describing a sequence; it is not evidence that a visual prompt recovers a soundtrack. Sound production has its own inputs, controls and quality checks.
The distinction also protects the quality of the final brief. A generator asked to obey invented dialogue, uncertain lip movement and a complex camera move at once may have an unnecessarily overloaded task. Keep the visual sequence usable even if the final audio is prepared later.
Can I infer the mood of the music from the pictures?
You can suggest a fitting mood, but label it as a creative interpretation. The original track could have been intentionally contradictory.
What should the returned prompt say when audio is unavailable?
Use a direct note such as original audio not established, followed by the visual instructions. Start in Reverse Prompt and keep any later soundtrack brief separate from the factual analysis.