Give each visible subject a stable descriptive label, then track actions separately within each shot. “The couple move toward the table” is ambiguous when only one person walks. Use appearance and screen position as temporary anchors, and re-establish those anchors after a cut instead of assuming left and right remain fixed.
Use labels that the footage can support
“Person in the ivory jacket” and “person in the dark waistcoat” are more useful than invented names or inferred relationships. Do not infer occupation, ethnicity or private identity from appearance. If the video itself establishes names and you are authorized to use them, keep the naming consistent with that evidence.
Screen position is useful but not permanent. A person on the left in a wide shot can be on the right after a reverse angle. Combine an appearance anchor with position in the current shot. If both people wear similar clothing, use another visible detail or note that identification is uncertain.
Make a subject-action grid
| Moment | Ivory jacket | Dark waistcoat | Shared space |
|---|---|---|---|
| Start | Stands beside the table | Faces the doorway | Closed box on table |
| Middle | Looks toward partner | Takes one step inward | Box remains untouched |
| End | Offers an open hand | Stops and looks down | No object exchange yet |
The grid prevents a broad summary from inventing synchronized movement. It also makes it easier to spot a missing reaction. If the source cuts away during an exchange, record the visible before and after states rather than inventing an unseen hand action.
Turn the grid into a local prompt
The person in the ivory jacket remains beside the table and turns their eyes toward the approaching partner. The person in the dark waistcoat enters from the doorway, takes one step and stops. At the end, the ivory-jacketed person offers an empty hand. The closed box stays on the table throughout.
This is an illustrative reconstruction, not an observation of an uploaded user clip. With your own footage, check the prompt against the actual frames. Distinguish an eye movement from a head turn and a held pose from an action the analyzer merely expects to happen.
Review cuts before transferring the description
Use Reverse Prompt to obtain a proposed breakdown, then verify which subject acts in each interval. Remove collective verbs when they conceal different behavior. Keep the resulting brief short enough that the destination generation tool can assign actions without several competing labels.
Runway's official prompting guide includes positional ways to describe multiple subjects. Its usefulness here is clarity of reference, not a guarantee of identity tracking by every model.
Can I label people A and B?
Yes in a planning note, but connect each label to visible evidence. An unexplained letter may mean nothing to a generator, especially if no matching reference labels are supported.
What if a person is hidden behind someone else?
Mark the action as unknown during the occlusion. Resume description when visible evidence returns. A continuous story does not justify filling every gap with a guessed movement.