成片拆解动作参考视频驱动的双人自拍整活:闺蜜手一挡一扫,前面的她在咧嘴笑和面无表情间来回切换,最后两人笑崩。
结构:两人板着脸 → 手扫过就变脸,反复切换 → 一起笑崩
看完成片再对照:这条片子靠什么在几秒内抓住人。点时间可跳到成片对应位置。
第 0 秒是手机前置自拍:戴橙色星星帽、扎双麻花辫的女孩在前,黑长发穿灰西装的闺蜜贴在她右后方,两人都面无表情看镜头。
约 0.5s 起,后面的闺蜜两只手一上一下交替从前面女孩脸上扫过;每扫一下,前面的她就在咧嘴大笑和闭嘴面无表情之间切换一次,闺蜜全程板着脸,一直持续到约 8.5s。
约 9s 闺蜜放下手,两人同时笑崩,张大嘴、眯眼、肩膀抖动,笑着结束。
动作和表情节奏全部交给一段真人动作参考视频,提示词只负责换人、换场景和补最后 1 秒的笑场。
别急着上视频模型。这个片子的秘密是:先把角色和场景锁死,再一次生成 1 个镜头。
image_1、image_2 是两个角色的设定图,每张都要有一张大尺寸脸部特写,全身图上的灰色圆圈是遮挡标记、不能出现在成片里;image_3 是房间图,决定卧室的结构、陈设、配色和光线。作者没有公开这 3 张图,也没有给出图提示词,需要自己准备。
前 9 秒的手部动作、表情切换时机、点头和两人站位,全部照搬 video_1。作者在回复里放出了他用的动作视频(一段两人自拍变脸的真人梗视频,16:9 带黑边),提示词里要求成片铺满画面、不要保留参考视频的黑边,也不要继承参考视频里人物的长相和衣服。
作者特别提醒:这段提示词是按他自己的角色和房间写的,不能直接照搬,要先丢给 GPT,让它按你的角色和场景改写。改好后连同 3 张图和动作视频一起交给视频模型。作者没有说明用的是哪个视频模型。
方便你对照成片理解每一镜在干什么。
直接全选复制到 Seedance。@Image 1–0 必须对应上面 REF 编号。
10s · 1:1 · image_1–3 + video_1 · 英文完整提示词(作者回复长帖)
SCENE CONTEXT A playful selfie video of two friends inside the colorful attic bedroom in <<<image_3>>>. Recreate the hand-sweep and expression-switching gag from <<<video_1>>> for the first nine seconds. During the final second, both break into spontaneous, wide-mouthed laughter. ACTIVE REFERENCES / REFERENCE USAGE <<<image_1>>> defines the foreground character’s identity, hair, body proportions, and complete outfit: orange star cap, twin braids, white keyhole crop top, cream jacket worn off the shoulders, and matching accessories. <<<image_2>>> defines the rear character’s identity, hair, body proportions, and complete outfit: long black hair with bangs, gray suit, white blouse, and black choker with an orange bead. Use the large facial portrait on each sheet for facial identity. The gray circles on the full-body views are reference masks and must not appear. <<<image_3>>> defines the room’s architecture, furnishings, colors, and lighting. Adapt its viewing angle to the selfie camera. <<<video_1>>> controls the first nine seconds of hand choreography, facial-expression timing, head movements, and relative character placement. Its performers’ appearances and clothing do not transfer. FIRST FRAME The video begins directly through the phone’s front-facing camera. <<<image_1>>> holds the phone at arm’s length, approximately at eye level, with her holding hand and phone outside the frame. Her face and upper chest fill the foreground. <<<image_2>>> is close behind her, slightly offset toward screen-right, with her face clearly visible beside <<<image_1>>>’s head. Both initially look into the lens with straight faces. WORLD AND SPATIAL BLOCKING They are positioned on the pink rug near the foot of the bed, with the blue photo-covered wall behind them. Portions of the pastel sloped ceiling and warm string lights remain visible above their heads; the bed and bright curtained window appear toward screen-right. Keep the room consistent with <<<image_3>>>, allowing the tight selfie framing to crop most furniture. <<<image_1>>> holds the phone throughout. <<<image_2>>> has both hands free to perform the gesture around <<<image_1>>>’s face. Maintain their front-to-back arrangement. SHOT FORMAT One continuous 10-second selfie take. The selfie image fills the output frame without the reference video’s black side padding. No cuts or external views of them filming. OPTICS AND CAMERA Natural front-camera perspective at arm’s length, with both faces readable and enough space above <<<image_1>>>’s cap for the hand movements. Keep the camera almost stationary during the gag, matching the reference’s stable composition with only slight natural hand drift. During the final laughter, allow a small, believable wobble from <<<image_1>>>’s shaking shoulders while keeping both faces in frame. No zoom, orbit, or dramatic reframing. ACTION / PERFORMANCE TIMING 0.0–0.7s: Match the reference’s brief neutral opening and hand preparation. <<<image_1>>> keeps a straight face. <<<image_2>>> raises one open hand above the cap and positions the other below <<<image_1>>>’s face. 0.7–9.0s: <<<image_2>>> reproduces the reference’s rapid alternating hand sweeps, exchanging the upper and lower hand positions and briefly obscuring <<<image_1>>>’s face as they pass. <<<image_1>>> switches between a broad toothy grin and a closed-mouth neutral expression at the corresponding moments in <<<video_1>>>. Preserve the original sweep directions, pace, short expression holds, blinks, and small head dips. <<<image_2>>> remains deliberately deadpan. Use the reference’s precise motion timing rather than inventing additional gestures. 9.0–10.0s: The gag breaks. <<<image_2>>> stops sweeping and lowers her hands clear of both faces. Both simultaneously burst into broad, open-mouthed laughter: cheeks lift, eyes crinkle, and shoulders bounce naturally. <<<image_2>>> leans slightly closer beside <<<image_1>>> so both laughing faces remain visible. <<<image_1>>> keeps holding the phone. End while both are still laughing, without a freeze or posed finish. PHYSICS AND CONTINUITY Maintain each character’s face, hairstyle, clothing, and accessories through all hand occlusions. Expressions change naturally without identity morphing. <<<image_2>>>’s hands remain anatomically connected to her arms. <<<image_1>>>’s orange cap stays securely on, her braids remain intact, and her cream jacket stays off her shoulders. Keep the hand choreography clear of the cap brim. Preserve natural hair and fabric movement. LIGHTING Use the bedroom’s soft window daylight from screen-right, balanced with the warm glow of the string lights. Keep both faces properly exposed and the pastel room recognizable. Lighting and exposure remain stable. AUDIO No added dialogue or narration. During the final second, both characters’ natural laughter begins in sync with their visible expressions, over quiet room ambience. LOCAL CONSTRAINTS Exactly two characters. Preserve the reference choreography’s screen direction without reversing it. Keep the phone outside its own camera view. No mirror shot, extra hands, face masks, captions, wardrobe changes, or unrelated action.
当前是单页静态站:视频 + 指引 + 可复制提示词。后面可以做成多教程列表、账号、收藏提示词、一键复制、手机竖屏版等。