Guides
ArtCraft Tutorial: From Prompt to Finished Shot
Updated 2026-10-11
ArtCraft's core idea is that generation is the last step, not the first. This tutorial walks the full workflow - prompt, composition, scene blocking, generation and refinement - the way the tool is designed to be used.
Step 1 - Write a prompt that describes the shot, not the vibe
Short, concrete prompts consistently beat long atmospheric ones. Describe the subject, its action, the camera and the light: "a barista pulling an espresso shot, low angle, warm morning light through the window" gives a model far more to work with than a paragraph of adjectives.
ArtCraft still lets you start from a prompt alone for quick exploration - but the point of the tool is that you do not stop there.
Step 2 - Compose on the 2D canvas
In the 2D workspace you arrange real elements: imported images, cutouts with backgrounds removed, backdrops and simple drawings. This is where you decide what is in frame and roughly where it sits.
- Import reference images or your own assets.
- Use background removal to cut a subject out cleanly.
- Layer backdrops, foreground props and subjects.
- Sketch guides if you want the model to follow a layout.
Step 3 - Block the scene in 3D (optional but powerful)
The signature feature: stage your shot in 3D before rendering. Turn an image into a 3D mesh, place objects with real depth, pose a character mannequin, and set the camera. For image-to-location work, you establish one consistent space and film multiple shots in it - so things do not wander between frames the way they do with prompt-only tools.
- Image-to-3D: convert reference images into positionable meshes.
- Character posing: set body pose and camera before generating.
- Scene blocking / kitbashing: combine asset kits to control angles and depth.
- Image-to-location: keep one environment consistent across shots.
Step 4 - Generate with the right model
Pick a model per task rather than one model for everything. Image models (Nano Banana, GPT Image, Seedream, Flux) differ in style and text handling; video models (Seedance, Kling, Sora, Veo) differ in motion quality, length and cost. Start with a fast/cheap option to iterate, then switch to a higher-fidelity model for the final render.
Step 5 - Refine locally instead of re-rolling the whole shot
The biggest workflow win: when a render is 90% right, fix the 10% in place. Re-run a masked region to repair hands, eyes or logos; adjust the canvas or scene and regenerate; keep camera and lighting stable across a set of related outputs. This is the difference between a slot machine and a tool.
Prompt patterns that work well here
- Subject + action + camera: "a violinist mid-bow, over-the-shoulder angle, shallow depth of field".
- Composition + style: "flat-design poster of a lighthouse, bold shapes, two-color palette".
- Reference-led: upload the product photo, then "place it on a stone plinth, soft studio light".
- Video: "slow push-in on the character, dust in the air, golden hour" - plus a blocked camera for precision.
Try it in your browser
Edit an image
No install, first generation free.