成片拆解
来源 · X

绫波丽双次吐舌表情深度驱动 · MiniMax H3
15 秒真人风短片,怎么做出来的?

MiniMax H3 表情驱动短片:一张参考图锁定绫波丽外观,深度动画带动她两次俏皮吐舌。

套路 · 角色表演15秒 · 表情驱动16:9 横屏MiniMax H3ReferenceToVideo深度动画驱动双次吐舌

钩子在哪

时间为抽帧估计

结构:固定机位 · 两次吐舌表情

看完成片再对照:这条片子靠什么在几秒内抓住人。点时间可跳到成片对应位置。

开场钩子

白墙前的自拍式正面近景,黑长直女孩穿浅蓝针织衫直视镜头,约 3s 就眨眼吐舌——第一个表情来得很快。

过程怎么推进

吐舌后闭眼说话,约 7s 抬手指戳脸颊,约 8–9s 手指点脸再吐一次舌,比第一次更夸张。

结尾怎么收

约 11s 手比「一点点」,之后放松微笑对镜头说话,没有额外剪辑。

跟做照抄这一点

机位和构图全程不动,只靠脸部表情撑住:两次吐舌卡在约 3s 和约 8.5s。

三步搞懂

别急着上视频模型。这个片子的秘密是:先把角色和场景锁死,再一次生成 1 个镜头。

1

先锁定身份与环境,只让 Picture 1 决定外观

提示词的 subject_definitions 把 <Subject 1>(绫波丽:同一张脸、自然肤质、深色眼睛、黑长直齐刘海、浅蓝圆领针织衫、银项链)和 <Subject 2>(白色纹理墙与沙发边)都绑定到 <Picture 1>,并声明 Picture 1 决定身份、服装、背景、光线、构图和首帧。实际复刻时先准备一张稳定的正面中近景 Picture 1 作为唯一外观基准。注意:作者没有公开这张 Picture 1,本页 refs 只是成片截帧,不能当作原始参考图。

2

把动作和外观分层:Video 1 只提供动作与时间

<Video 1> 是带面部运动曲线和舌头轮廓的深度诊断动画,提示词明确“只把它读作动作及其时间的指引;所有可见表面、颜色、面部比例和光照都只来自 Picture 1”。retention_analysis 进一步要求保持原有下颌宽度、脸颊体积和下巴长度,并把动作转译为自然皮肤与针织面料。这是避免模型把灰度深度图、彩色曲线或点直接渲染进成片的关键。

3

用时间码写两次吐舌,最后写负向连续性

detailed_description 用时间码约束:0.00–3.20 秒舌头在口内;约 3.20 秒伸出、约 4.07 秒收回;约 8.81–10.68 秒第二次伸出、短暂停留并收回;10.68–15.00 秒按原节奏继续其余表情。强调“Two tongue cycles only”,禁止循环、变速、重置和结尾冻结;固定机位与脸部尺度,不改毛衣和墙面;禁止灰度深度图、彩色面部曲线、点、遮罩、文字、图形、断开的舌头和变形手臂。原声音频由作者外部添加,提示词要求不生成对白。

步骤一 · 做出这 5 张参考图

推荐用 GPT Image / 同等图模,比例 4:3。每张点开即可复制完整英文提示词。

1

开场中近景

t≈1s · 成片截帧

第 1 步
开场中近景
▸查看完整出图提示词(点击复制区)
复制
成片截帧(非作者参考图/非 Picture 1):固定机位、写实中近景,绫波丽黑长直齐刘海、浅蓝针织衫,白色纹理墙与沙发边,舌头仍在口内,只有细微眼神与头部动作。从作者发布的彩色成片抽帧,仅作跟随拆解;作者未公开静态 Picture 1。
2

第一次吐舌

t≈3.5–4s · 成片截帧

第 2 步
第一次吐舌
▸查看完整出图提示词(点击复制区)
复制
成片截帧(非作者参考图/非 Picture 1):第一轮吐舌(提示词约 3.20–4.07 秒):单眼眨眼、粉色舌头伸过下唇,头部微侧,脸型与服装保持不变。从作者发布的彩色成片抽帧,仅作跟随拆解;作者未公开静态 Picture 1。
3

两次吐舌之间

t≈7s · 成片截帧

第 3 步
两次吐舌之间
▸查看完整出图提示词(点击复制区)
复制
成片截帧(非作者参考图/非 Picture 1):两轮吐舌之间的过渡表情:舌头已收回,嘴唇微张,正视镜头,机位与脸部尺度与开场一致。从作者发布的彩色成片抽帧,仅作跟随拆解;作者未公开静态 Picture 1。
4

第二轮吐舌区间

t≈10s · 成片截帧

第 4 步
第二轮吐舌区间
▸查看完整出图提示词(点击复制区)
复制
成片截帧(非作者参考图/非 Picture 1):处于提示词第二轮吐舌区间(约 8.81–10.68 秒)附近:嘴部张开、表情变化,肩部和针织袖子保持连续。从作者发布的彩色成片抽帧,仅作跟随拆解;作者未公开静态 Picture 1。
5

收尾细微表情

t≈13s · 成片截帧

第 5 步
收尾细微表情
▸查看完整出图提示词(点击复制区)
复制
成片截帧(非作者参考图/非 Picture 1):10.68 秒之后的收尾段:闭眼、嘴部细微动作,继续原节奏的眼、嘴、头部表情,无循环或结尾冻结。从作者发布的彩色成片抽帧,仅作跟随拆解;作者未公开静态 Picture 1。

步骤二 · 1 镜头故事板(人话版)

方便你对照成片理解每一镜在干什么。

  1. Shot 1 — 完整 15 秒单镜头(固定机位、平视、16:9 写实中近景):从 Picture 1 首帧开始。0.00–3.20 秒完成最初的眼部、嘴唇和细微头部动作,舌头在口内;约 3.20 秒嘴唇分开,自然粉色舌头伸过下唇,约 4.07 秒收回;中间继续表情与小手势;约 8.81–10.68 秒第二次伸舌、短暂停留并收回;10.68–15.00 秒按原节奏继续其余眼、嘴、头和手部动作。全程同一张脸、黑发、浅蓝毛衣、项链和白墙,无对白、无剪切、无循环或结尾冻结。
关键约束:Picture 1 是唯一外观来源(脸/发型/浅蓝针织衫/银项链/白墙/光线/首帧);Video 1 深度诊断动画只读动作与时间,不得渲染灰度深度图、彩色面部曲线、点、遮罩、文字或图形;仅两次吐舌(约3.20–4.07秒、约8.81–10.68秒),无循环/变速/重置/结尾冻结;固定机位与原脸部尺度;保持下颌宽度、脸颊体积、下巴长度与连续针织袖子;禁止断开的舌头和变形手臂;原声音频由作者外部添加,不生成对白。缺口:作者未公开静态 Picture 1,refs 全部为成片截帧(非参考图);Video 1 深度/面部曲线动画随帖发布但为视频而非静态参考图,未在本页收录;未发现 quoted post 或额外提示词;demo-web.mp4 为网页压缩版,音频未做转录。

步骤三 · 完整视频提示词

直接全选复制到 Seedance。@Image 1–5 必须对应上面 REF 01–05。

GEN 01

Rei 双次吐舌 · ReferenceToVideo 提示词

15s · 16:9 · MiniMax H3 ReferenceToVideo · Picture 1 + Video 1 · 英文完整提示词

复制
subject_definitions:
<Subject 1> is Rei in <Picture 1>: the exact same face, natural skin, dark eyes, straight long black hair with blunt bangs, light-blue crewneck knitted sweater and silver necklace.
<Subject 2> is the white textured wall and sofa edge in <Picture 1>.
<Picture 1> defines identity, clothing, background, lighting, framing and the opening image.
<Video 1> is a diagnostic depth animation with facial motion curves and a tongue silhouette. Read it only as a guide to movements and their timing; all visible surfaces, colors, facial proportions and illumination come exclusively from <Picture 1>.

summary:
[reference generation] A continuous fifteen-second photorealistic medium close-up of <Subject 1>, fixed eye-level camera, landscape 16:9. Keep the same face, black hair, light-blue sweater, necklace and white wall throughout. She makes two brief playful tongue-out gestures at the original reference times, retracts her tongue after each and continues subtle facial expressions. No speech or cuts.

retention_analysis:
Fully preserve <Picture 1> appearance and environment. Change only facial expression and small head movements. Preserve natural shoulders and continuous knitted sleeves. Translate the motion into natural skin and knitted cloth. The visible result is a normally illuminated color photograph of Rei, with her original jaw width, cheek volume and chin length.

detailed_description:
[Shot 1]
Begin from <Picture 1>, then smoothly perform the expression. From 0.00 to 3.20 seconds, perform the initial eye, lip and small head movements with the tongue inside the mouth. Around 3.20 seconds the lips part and a naturally pink tongue extends over the lower lip; it retracts around 4.07 seconds. Continue the intervening expressions and small gestures. Around 8.81 to 10.68 seconds, perform the second tongue extension, brief hold and retraction. From 10.68 to 15.00 seconds, continue the remaining eye, mouth, head and hand gestures at the original pace. Two tongue cycles only, naturally connected inside the mouth. No looping, speed change, reset or forced freeze at the end. Keep the face recognizable and skin realistic. Keep the camera fixed and the original face scale. Do not change the light-blue sweater or white wall. No gray depth-map rendering, colored facial curves, dots, masks, text or graphics. No detached tongue or distorted arms. Original audio is added separately; do not generate speech.

这是 Demo

当前是单页静态站:视频 + 指引 + 可复制提示词。后面可以做成多教程列表、账号、收藏提示词、一键复制、手机竖屏版等。

返回首页,看更多教程