Skip to content

Guides

ArtCraft Tutorial: From Prompt to Finished Shot

Updated 2026-10-11

ArtCraft's core idea is that generation is the last step, not the first. This tutorial walks the full workflow - prompt, composition, scene blocking, generation and refinement - the way the tool is designed to be used.

Step 1 - Write a prompt that describes the shot, not the vibe

Short, concrete prompts consistently beat long atmospheric ones. Describe the subject, its action, the camera and the light: "a barista pulling an espresso shot, low angle, warm morning light through the window" gives a model far more to work with than a paragraph of adjectives.

ArtCraft still lets you start from a prompt alone for quick exploration - but the point of the tool is that you do not stop there.

Step 2 - Compose on the 2D canvas

In the 2D workspace you arrange real elements: imported images, cutouts with backgrounds removed, backdrops and simple drawings. This is where you decide what is in frame and roughly where it sits.

  • Import reference images or your own assets.
  • Use background removal to cut a subject out cleanly.
  • Layer backdrops, foreground props and subjects.
  • Sketch guides if you want the model to follow a layout.

Step 3 - Block the scene in 3D (optional but powerful)

The signature feature: stage your shot in 3D before rendering. Turn an image into a 3D mesh, place objects with real depth, pose a character mannequin, and set the camera. For image-to-location work, you establish one consistent space and film multiple shots in it - so things do not wander between frames the way they do with prompt-only tools.

  • Image-to-3D: convert reference images into positionable meshes.
  • Character posing: set body pose and camera before generating.
  • Scene blocking / kitbashing: combine asset kits to control angles and depth.
  • Image-to-location: keep one environment consistent across shots.

Step 4 - Generate with the right model

Pick a model per task rather than one model for everything. Image models (Nano Banana, GPT Image, Seedream, Flux) differ in style and text handling; video models (Seedance, Kling, Sora, Veo) differ in motion quality, length and cost. Start with a fast/cheap option to iterate, then switch to a higher-fidelity model for the final render.

Step 5 - Refine locally instead of re-rolling the whole shot

The biggest workflow win: when a render is 90% right, fix the 10% in place. Re-run a masked region to repair hands, eyes or logos; adjust the canvas or scene and regenerate; keep camera and lighting stable across a set of related outputs. This is the difference between a slot machine and a tool.

Prompt patterns that work well here

  • Subject + action + camera: "a violinist mid-bow, over-the-shoulder angle, shallow depth of field".
  • Composition + style: "flat-design poster of a lighthouse, bold shapes, two-color palette".
  • Reference-led: upload the product photo, then "place it on a stone plinth, soft studio light".
  • Video: "slow push-in on the character, dust in the air, golden hour" - plus a blocked camera for precision.

Try it in your browser

Edit an image

No install, first generation free.

Open tool

Common questions

Do I have to use the 3D scene tools?
No - plain text-to-image and text-to-video work fine. The 3D blocking is there for when you need repeatable, controllable results.
Why do my generations drift between shots?
Prompt-only generation has no memory of space. Use image-to-location or a blocked 3D scene so the environment and camera stay consistent across a set.