成片拆解
来源 · X

手一扫就变脸 · 双人自拍整活
10 秒自拍整活短片,怎么做出来的?

动作参考视频驱动的双人自拍整活:闺蜜手一挡一扫,前面的她在咧嘴笑和面无表情间来回切换,最后两人笑崩。

套路 · 角色表演套路 · 手机POV·Vlog10秒 · 自拍整活1:1 方屏3 张参考图 + 动作参考视频一镜到底 · 手机前置视角变脸梗 · 双人

钩子在哪

时间为抽帧估计

结构:两人板着脸 → 手扫过就变脸,反复切换 → 一起笑崩

看完成片再对照:这条片子靠什么在几秒内抓住人。点时间可跳到成片对应位置。

开场钩子

第 0 秒是手机前置自拍:戴橙色星星帽、扎双麻花辫的女孩在前,黑长发穿灰西装的闺蜜贴在她右后方,两人都面无表情看镜头。

过程怎么推进

约 0.5s 起,后面的闺蜜两只手一上一下交替从前面女孩脸上扫过;每扫一下,前面的她就在咧嘴大笑和闭嘴面无表情之间切换一次,闺蜜全程板着脸,一直持续到约 8.5s。

结尾怎么收

约 9s 闺蜜放下手,两人同时笑崩,张大嘴、眯眼、肩膀抖动,笑着结束。

跟做照抄这一点

动作和表情节奏全部交给一段真人动作参考视频,提示词只负责换人、换场景和补最后 1 秒的笑场。

三步搞懂

别急着上视频模型。这个片子的秘密是:先把角色和场景锁死,再一次生成 1 个镜头。

1

第一步:准备 3 张参考图(image_1–3)

image_1、image_2 是两个角色的设定图,每张都要有一张大尺寸脸部特写,全身图上的灰色圆圈是遮挡标记、不能出现在成片里;image_3 是房间图,决定卧室的结构、陈设、配色和光线。作者没有公开这 3 张图,也没有给出图提示词,需要自己准备。

2

第二步:准备动作参考视频(video_1)

前 9 秒的手部动作、表情切换时机、点头和两人站位,全部照搬 video_1。作者在回复里放出了他用的动作视频(一段两人自拍变脸的真人梗视频,16:9 带黑边),提示词里要求成片铺满画面、不要保留参考视频的黑边,也不要继承参考视频里人物的长相和衣服。

3

第三步:先让 GPT 按你的角色改写,再贴进视频模型

作者特别提醒:这段提示词是按他自己的角色和房间写的,不能直接照搬,要先丢给 GPT,让它按你的角色和场景改写。改好后连同 3 张图和动作视频一起交给视频模型。作者没有说明用的是哪个视频模型。

步骤二 · 1 镜头故事板(人话版)

方便你对照成片理解每一镜在干什么。

  1. Shot 1 — 0–0.7s 前置自拍,两人面无表情看镜头;后面的闺蜜一只手举到帽子上方,另一只手放在前面女孩脸下方。
  2. Shot 2 — 0.7–9s 闺蜜两只手上下交替扫过前面女孩的脸,前面的她在咧嘴大笑和闭嘴面无表情之间来回切换,闺蜜始终板着脸。
  3. Shot 3 — 9–10s 闺蜜放下手,两人同时大笑,肩膀抖动,画面轻微晃动,笑着结束。
关键约束:只能有两个人;前面的女孩全程拿着手机,手和手机都不能入镜;一镜到底,不剪辑、不变焦、不绕拍;手扫过时人物的脸、发型、衣服和配饰不能变,帽子不能掉,外套保持露肩;保持参考动作的方向,不要左右镜像;不要镜子画面、多余的手、字幕和换装;最后 1 秒两人的笑声和表情同步。与成片不符:后面闺蜜的手腕上戴着一块表,提示词里没写,像是从动作参考视频带过来的。缺口:作者没公开 image_1–3 这 3 张参考图,也没给出图提示词;成片右上角有一个像素风「G」字水印,看不出是哪个平台;动作参考视频是真人素材,只留在本地,没有放进教程;音频没有核对。

步骤三 · 完整视频提示词

直接全选复制到 Seedance。@Image 1–0 必须对应上面 REF 编号。

GEN 01

双人自拍变脸 · 视频提示词

10s · 1:1 · image_1–3 + video_1 · 英文完整提示词(作者回复长帖)

复制
SCENE CONTEXT
A playful selfie video of two friends inside the colorful attic bedroom in <<<image_3>>>. Recreate the hand-sweep and expression-switching gag from <<<video_1>>> for the first nine seconds. During the final second, both break into spontaneous, wide-mouthed laughter.

ACTIVE REFERENCES / REFERENCE USAGE
<<<image_1>>> defines the foreground character’s identity, hair, body proportions, and complete outfit: orange star cap, twin braids, white keyhole crop top, cream jacket worn off the shoulders, and matching accessories.
<<<image_2>>> defines the rear character’s identity, hair, body proportions, and complete outfit: long black hair with bangs, gray suit, white blouse, and black choker with an orange bead.
Use the large facial portrait on each sheet for facial identity. The gray circles on the full-body views are reference masks and must not appear.
<<<image_3>>> defines the room’s architecture, furnishings, colors, and lighting. Adapt its viewing angle to the selfie camera.
<<<video_1>>> controls the first nine seconds of hand choreography, facial-expression timing, head movements, and relative character placement. Its performers’ appearances and clothing do not transfer.

FIRST FRAME
The video begins directly through the phone’s front-facing camera. <<<image_1>>> holds the phone at arm’s length, approximately at eye level, with her holding hand and phone outside the frame. Her face and upper chest fill the foreground. <<<image_2>>> is close behind her, slightly offset toward screen-right, with her face clearly visible beside <<<image_1>>>’s head. Both initially look into the lens with straight faces.

WORLD AND SPATIAL BLOCKING
They are positioned on the pink rug near the foot of the bed, with the blue photo-covered wall behind them. Portions of the pastel sloped ceiling and warm string lights remain visible above their heads; the bed and bright curtained window appear toward screen-right.
Keep the room consistent with <<<image_3>>>, allowing the tight selfie framing to crop most furniture.
<<<image_1>>> holds the phone throughout. <<<image_2>>> has both hands free to perform the gesture around <<<image_1>>>’s face. Maintain their front-to-back arrangement.

SHOT FORMAT
One continuous 10-second selfie take. The selfie image fills the output frame without the reference video’s black side padding. No cuts or external views of them filming.

OPTICS AND CAMERA
Natural front-camera perspective at arm’s length, with both faces readable and enough space above <<<image_1>>>’s cap for the hand movements.
Keep the camera almost stationary during the gag, matching the reference’s stable composition with only slight natural hand drift. During the final laughter, allow a small, believable wobble from <<<image_1>>>’s shaking shoulders while keeping both faces in frame. No zoom, orbit, or dramatic reframing.

ACTION / PERFORMANCE TIMING
0.0–0.7s:
Match the reference’s brief neutral opening and hand preparation. <<<image_1>>> keeps a straight face. <<<image_2>>> raises one open hand above the cap and positions the other below <<<image_1>>>’s face.

0.7–9.0s:
<<<image_2>>> reproduces the reference’s rapid alternating hand sweeps, exchanging the upper and lower hand positions and briefly obscuring <<<image_1>>>’s face as they pass.
<<<image_1>>> switches between a broad toothy grin and a closed-mouth neutral expression at the corresponding moments in <<<video_1>>>. Preserve the original sweep directions, pace, short expression holds, blinks, and small head dips. <<<image_2>>> remains deliberately deadpan.
Use the reference’s precise motion timing rather than inventing additional gestures.

9.0–10.0s:
The gag breaks. <<<image_2>>> stops sweeping and lowers her hands clear of both faces. Both simultaneously burst into broad, open-mouthed laughter: cheeks lift, eyes crinkle, and shoulders bounce naturally. <<<image_2>>> leans slightly closer beside <<<image_1>>> so both laughing faces remain visible. <<<image_1>>> keeps holding the phone. End while both are still laughing, without a freeze or posed finish.

PHYSICS AND CONTINUITY
Maintain each character’s face, hairstyle, clothing, and accessories through all hand occlusions. Expressions change naturally without identity morphing. <<<image_2>>>’s hands remain anatomically connected to her arms.
<<<image_1>>>’s orange cap stays securely on, her braids remain intact, and her cream jacket stays off her shoulders. Keep the hand choreography clear of the cap brim. Preserve natural hair and fabric movement.

LIGHTING
Use the bedroom’s soft window daylight from screen-right, balanced with the warm glow of the string lights. Keep both faces properly exposed and the pastel room recognizable. Lighting and exposure remain stable.

AUDIO
No added dialogue or narration. During the final second, both characters’ natural laughter begins in sync with their visible expressions, over quiet room ambience.

LOCAL CONSTRAINTS
Exactly two characters. Preserve the reference choreography’s screen direction without reversing it. Keep the phone outside its own camera view. No mirror shot, extra hands, face masks, captions, wardrobe changes, or unrelated action.

这是 Demo

当前是单页静态站:视频 + 指引 + 可复制提示词。后面可以做成多教程列表、账号、收藏提示词、一键复制、手机竖屏版等。

返回首页,看更多教程