ArtArch 新闻中心
AI Video TechniquesAI Image Prompt Structure: An 8-Part Framework Tested
Learn an eight-part AI image prompt structure for controlling subject, scene, action, emotion, composition, lighting, style, and quality.

Provide a character reference and a neutral location reference to create a visually focused rainy-night portrait with deliberate action, emotion, composition, and light. This AI image prompt structure organizes the direction into eight parts: subject, scene, action, emotion, composition, lighting, style, and quality constraints. AVT-014 tested the framework in a café shot and compared it with a prompt built from broad phrases such as cinematic, realistic, detailed, moody, and beautiful.
Both versions used the same short-haired woman in a dark green sweater, the same corner café beside a rain-covered window, and Seedance 2.0 Fast at 1280 × 720. The control created an attractive wide café scene. The structured version moved closer, kept both hands and the cup prominent, gave the woman's gaze a clear path, and divided the warm interior light from the cold blue street outside.
The framework helps because each module answers one production question. The creator can see what belongs in the frame, what the subject does, how the emotion becomes visible, where attention should land, and which details need protection.
Open the AVT-014 structured-prompt experiment on ArtArch.

The character reference fixes the woman's short brown bob, dark green sweater, black trousers, and natural proportions before the A/B comparison.

The empty location reference gives both prompts the same corner table, pendant light, wet window, brick wall, and street reflections.
What the AVT-014 A/B test showed
Both nodes were configured for eight seconds, while the completed files available for this review run for approximately 5.04 seconds. The comparison therefore uses frames from 0.2 seconds through 4.6 seconds. The review covers subject priority, action order, emotional evidence, camera distance, screen placement, light logic, color separation, and object visibility.
| Observation | A: descriptive mood words | B: eight-part structure |
|---|---|---|
| Camera distance | Wide view showing the table, chairs, full window, street, and figure | Side medium-close view centered on face, hands, cup, and window edge |
| Subject placement | Figure remains small beside a large area of rainy glass | Figure occupies the right side and carries most of the visual weight |
| Action | Holds the cup, looks around, then drinks | Holds the cup with both hands, breathes near the rim, then looks toward the window |
| Emotion | Mood comes mainly from the rainy exterior and solitary setting | Fatigue appears through lowered gaze, contained posture, and the later window look |
| Lighting | Warm table and skin against a broad blue street view | Warm overhead side light models the face; cool exterior light outlines the profile |
| Visual focus | Environment and person share attention | Eyes, hands, and cup form one readable attention triangle |
| Object visibility | Cup is clear during the opening and drinking beat, then leaves the final close-up | Cup shape and two-hand relationship remain visible across the sampled frames |
A group: descriptive mood words

At 0.2 seconds, the control begins as a wide establishing view in which the setting and person share attention.

Near 2.5 seconds, the camera has moved closer and the woman raises the cup to drink.

Near 4.6 seconds, the push-in prioritizes her face while the cup and hands have moved outside the frame.
<video controls playsinline preload="metadata" src="https://d3ohi5svx56rr5.cloudfront.net/artarchtask/ec52d9c3-dae0-40e1-baab-30130b1a41e3/url_0770e8c29662a76ce480aa93935ab86a.mp4"></video>
A group: broad mood words produce a polished rainy-night scene and leave the model to choose the changing visual priority.
B group: eight-part prompt structure

At 0.2 seconds, the structured prompt establishes the woman, both hands, cup, table, and rainy window in one readable relationship.

Near 2.5 seconds, the side view keeps the cup stable while her gaze moves away from it.

Near 4.6 seconds, the final sampled frame retains the same focal group and the window-directed gaze.
<video controls playsinline preload="metadata" src="https://d3ohi5svx56rr5.cloudfront.net/artarchtask/ec52d9c3-dae0-40e1-baab-30130b1a41e3/url_3a744c3a90b1caf07b57b92f59bf971a.mp4"></video>
B group: the eight modules keep the subject, hands, cup, gaze, light, and background working toward one visual event.
The control offers a polished establishing image. Its rain, reflections, blue night, warm table, and solitary customer communicate atmosphere immediately. The prompt leaves the model free to decide where the emotional story lives, so the setting carries more meaning at the opening and the face becomes dominant after the push-in.
The technique version gives each visual element a job. The woman is the subject. The window seat is the scene. Holding the cup and shifting her gaze form the clearest visible action. Tired eyes and contained posture carry the emotion. The medium-close side view places the eyes, hands, and cup at the center. Warm interior light and cool street light separate the two emotional spaces. The planned slow breath and shoulder tension are subtler than the grip and gaze change, so the evidence supports the broader performance direction more strongly than those exact micro-beats.
This is an E1 observation from one A/B pair with open video seeds. It shows a clearer hierarchy in this scene and provides a framework for repeat tests with other characters, locations, and models.
The eight modules of a structured prompt
1. Subject: define the visual anchor
Begin with the person, object, creature, or product that carries the image. Include only the stable traits needed for recognition.
A short-haired adult woman wearing a dark green knitted sweater.
The subject line creates a hierarchy before scenery and style enter the prompt. For a character, useful anchors include age range, hairstyle, wardrobe, body proportions, and one distinctive feature. For a product, use shape, material, color, and important design details.
2. Scene: locate the subject in a specific space and time
Name the place, time, and two or three spatial anchors.
She sits at the window seat of a corner café late at night.
Rain covers the glass, with distant streetlights beyond it.
Specific anchors help composition, depth, and lighting work together. The AVT-014 scene includes a wood table, window, rain texture, and distant streetlight. Each item supports the lonely waiting moment.
3. Action: choose one visible event
Describe a small action sequence with a beginning and endpoint.
She holds a plain coffee cup with both hands.
She exhales slowly near the rim, then moves her gaze from the cup toward the window.
Action gives the frame narrative direction. It also gives the hands, prop, gaze, and timing a shared purpose. A single continuous event is especially useful for short video generation and keyframes designed to lead into motion.
4. Emotion: translate feeling into evidence
Connect the intended emotion to observable signals.
Her eyes look tired.
Her shoulders draw in slightly.
Her breathing stays controlled.
The expression suggests that someone has failed to arrive.
“Sad” or “moody” names an interpretation. Eyes, mouth, posture, breath, grip, and gaze direction make that interpretation visible. In AVT-014, the control relies heavily on rain and solitude, while the structured version adds performance evidence inside the frame.
5. Composition: decide where attention lands
State shot size, angle, placement, and the primary focus.
Use a side medium-close view.
Place the woman on the right side of the frame.
Keep the eyes, both hands, and cup as the main focus.
Retain a simple section of the rainy window behind her.
Composition turns a list of objects into a reading order. In the technique result, the close relationship between face, fingers, cup, and window creates a clear path for the eye.
6. Lighting: describe a physical source and relationship
Name where light comes from, how it falls, and how sources interact.
A warm pendant light from the upper left creates soft side illumination.
Cool streetlight through the window adds a faint edge along her profile.
Keep the cup and hands readable within the warm pool of light.
The warm interior and cold exterior become part of the story. One source contains the character; the other marks the world she watches.
7. Style: set the overall visual language
Choose one main style and a small supporting treatment.
Observational cinematic realism.
Restrained saturation and subtle natural grain.
Style works after the subject, action, composition, and light have established the scene. It unifies the decisions already present.
8. Quality and constraints: protect the fragile details
End with a short list tied to the actual risks of the image.
Preserve natural skin texture, facial identity, finger anatomy, cup shape, and café layout.
Keep the visual focus on the woman, hands, and cup.
Prevent extra people, decorative clutter, cup deformation, and focus drift.
Quality becomes useful when it names what must stay stable. “Detailed” is broad; “stable cup rim, natural fingers, readable skin texture” provides inspectable targets.
A reusable AI image prompt structure
Subject:
[Who or what is the visual anchor? Include stable identity or product traits.]
Scene:
[Where and when is the subject? Name two or three spatial anchors.]
Action:
[What single visible event is happening? Give it an endpoint.]
Emotion:
[Which eyes, mouth, posture, breath, grip, or gaze signals carry the feeling?]
Composition:
[Choose shot size, camera angle, subject placement, and the primary focus.]
Lighting:
[Name the physical source, direction, softness, temperature, and relationship between sources.]
Style:
[Choose one main visual language and one supporting texture or finish.]
Quality and constraints:
[Protect identity, anatomy, prop shape, material texture, spatial layout, and visual priority.]
Write each section as a production decision. When the result drifts, revise the module responsible for that part of the image.
How the modules work together
The strongest prompts create agreement between modules.
In AVT-014, the subject and action require visible hands and a cup. The composition therefore moves to a medium-close view. The emotion requires tired eyes, contained breath, and a window gaze. The side angle makes those signals readable. The scene provides rain and distant streetlights. The lighting turns them into a cool rim around a warm interior portrait.
Each decision supports the same focal event: a woman waits with a cooling drink and realizes the meeting will remain empty.
Conflicts become easier to spot inside this structure. A full-body composition would weaken the tiny change in her eyes. A rapid camera move would compete with controlled breathing. A bright fashion palette would shift the emotional reading. A crowded background would divide attention from the face and cup.
Use the framework for different creation scenes
Character portraits
A filmmaker preparing a dramatic character portrait can define identity, location, one physical behavior, visible emotional signals, lens relationship, motivated light, and a short list of anatomy protections. The framework keeps wardrobe, posture, and setting connected to the same story beat.
Product lifestyle images
A brand creator can make the product the subject, define the real use environment, show one hand interaction, choose a composition that preserves the label or silhouette, and specify how the material responds to light. Quality constraints can protect product shape and surface detail.
Editorial fashion frames
A fashion team can anchor the model, garment, environment, pose transition, attitude, negative space, hard or soft light, and fabric behavior. The structure separates wardrobe identity from styling language and composition.
Travel and hospitality scenes
A travel creator can identify the traveler, location architecture, action, emotional response, environmental scale, time-of-day light, and local materials. A clear hierarchy keeps the person connected to the destination while preserving a strong sense of place.
Storyboard keyframes
A director can use each prompt as one beat: subject, location, action, performance, framing, light, style, and continuity locks. Repeating stable modules across several keyframes supports a coherent sequence while action and composition change with the story.
Diagnose a generated result by module
Subject drift
Strengthen stable identity, wardrobe, product shape, or reference roles. Remove competing descriptions that introduce a second visual anchor.
Weak action
Reduce the event to one visible sequence and name its endpoint. Connect hands, gaze, prop, and body movement to that action.
Generic emotion
Replace the feeling label with two or three performance signals. Choose signals that fit the shot size and camera angle.
Unclear composition
State where the subject sits in the frame, how large they appear, and which body part or object receives focus. Simplify background anchors.
Flat lighting
Give the illumination a real source, direction, softness, and temperature. Describe how it affects the face, hands, prop, and background separately.
Artificial texture
Name the materials that matter and how they respond to light. Preserve natural skin variation, fabric weave, wood grain, condensation, rain, ceramic glaze, or metal reflections as appropriate.
Unstable props or hands
Keep the interaction simple, state the grip or contact clearly, and protect anatomy and object shape in the final module. Review early, middle, and final frames.
Review the final image or clip
Start by removing the style words in your mind. You should still be able to identify the subject, location, action, emotional evidence, composition, and light logic.
Then check whether one visual anchor dominates. In AVT-014, the technique frame establishes the woman's face, hands, and cup as a single focal group. The rainy window supports the emotion without taking control of the composition.
Finally, compare the output with each module. Record partial success precisely: the action may work while gaze timing drifts, or the light may work while the cup shape changes. Module-level review leads to targeted revisions.
Frequently asked questions
What is the best order for an AI image prompt?
Start with subject, scene, and action; continue with visible emotion, composition, and lighting; finish with style plus quality and stability constraints.
How long should a structured AI image prompt be?
Use enough detail to make each of the eight production decisions clear. Keep one primary subject, one action, a small set of scene anchors, and a focused list of constraints.
How do I make an AI portrait feel less generic?
Give the person a specific task, translate emotion into small physical signals, choose a purposeful camera distance, use a believable light source, and preserve natural material and skin detail.
Can the same structure guide AI video prompts?
Yes. Keep the eight modules, then express the action and performance as a short ordered sequence. Add camera movement only when it supports the focal event.
继续探索
更多 AI Video Techniques 内容

Motivated Cut AI Video Continuity: Change the Shot for a Reason
Learn how a visible action trigger, clear shot-size change, and locked screen direction create more intentional same-scene AI video edits.

Deep Depth of Field AI Video: Keep Every Story Layer Readable
Learn how deep depth of field AI video prompts keep foreground tools, central action, and distant location cues readable in one layered shot.

Medium Depth of Field AI Video: Balance Subject and Setting
Learn how medium depth of field AI video prompts keep hands and active objects readable while preserving recognizable environmental context.

AI Video Character Reference Framing: Match Portraits to Landscape Shots
Prepare vertical character portraits for landscape AI video by matching aspect ratio, body scale, headroom, floor space, and movement room before generation.

