Record a continuous performance that makes the line, expression and pauses understandable together. Stabilize the phone, keep the relevant face and gestures visible, and review the recording before using it as a driving reference. A clean sentence in a text document cannot show the gaze change or hesitation you want the character to inherit.
Prepare one delivery, not an entire script
Choose a short event: a welcome, a question or a quiet response. Read it aloud until it sounds natural without becoming memorized shouting. Mark any name that needs a particular pronunciation and any pause that changes the meaning.
For a fictional proposal scene, the line might be a brief invitation to look at a letter, followed by a pause. Decide whether the performer looks toward the letter before speaking or after the invitation. Those alternatives carry different intentions. Record the version you actually want rather than expecting a generator to infer it.
Runway's dialogue workflow begins with recorded performances. This is a preparation method; it does not mean every video generator accepts acting footage.
Set up the phone around the performance
Use a stable support and a quiet place. Check the framing by recording a few seconds, including the widest planned gesture. A hand that enters only halfway through can be harder to interpret than one visible from the beginning. Keep the face large enough to read the gaze and mouth without requiring a tight crop that cuts off needed movement.
Act-Two's performance guidance calls for a visible face and an uninterrupted shot. Follow the selected renderer's own framing and duration requirements; a dialogue-performance setup should not be assumed suitable for full-body dancing.
Avoid distracting activity behind the performer. The source is easier to review when you can tell which movement belongs to the intended acting.
Review speech and movement together
| Review pass | What to listen or look for |
|---|---|
| Normal playback | Natural delivery and understandable emotion |
| Audio-only | Clear words, names and meaningful pauses |
| Muted playback | Gaze, posture and gesture convey the intended beat |
| Beginning and finish | Enough usable lead-in and a completed response |
Choose one take because it works as a performance, not because it has the highest resolution. A sharper recording with the wrong expression is still the wrong source. Keep the original and mark the useful interval in your notes rather than overwriting it with a tightly trimmed copy.
If two characters converse, record their roles separately and preserve the intended timing. Do not make one performance file responsible for deciding which character owns every line.
Keep the handoff explicit
Include the approved line, the source interval, the intended character and a small list of preserved traits in the generation brief. Specify whether the gesture matters or whether you mainly need facial delivery. The renderer's supported input mode determines what can actually transfer.
An Imagild film plan can organize these requirements. It does not establish that the current renderer supports this competitor's performance-capture or character-video gesture settings. Confirm those capabilities before a paid attempt.
Review the returned dialogue against the source at normal speed. Check that the pause remains meaningful, the mouth movement belongs to the correct speaker and the final reaction completes. If the tone is wrong, revise the performance or route deliberately; adding “emotional” repeatedly to the prompt is a poor substitute for a readable source.