Understanding Generative AI · Shoot First: The Craft
Single Shot · The Craft

Understanding Generative AI

What it actually does, and why that changes how you use it

Start here

Before any control below matters, know which kind of model you're using. It decides whether you'll ever touch these dials.

Commercial · managed
Kling, Seedance, Veo, Runway

The service wraps the model and runs it for you. You prompt; it handles the rest.

  • Prompt tuning, enhancement, and upscaling happen automatically
  • Reliable results with almost no setup
  • The controls in this sheet are handled behind the scenes
Best when: you want dependable results fast. Start here.
Open-weight · hands-on
WAN, LTX, and other local models

You run the model yourself, usually in ComfyUI, with every dial exposed.

  • Where Steps, CFG, Denoise, and ControlNet live
  • Where LoRAs, custom pipelines, and deep tinkering happen
  • Higher ceiling once you put the hours in
Best when: you want to fine-tune the result or build a repeatable pipeline. The rest of this sheet is your control panel.

Commercial wins on reliability out of the box. Open-weight wins on control and ceiling once you invest the time. Pick the one that matches how you want to work.

On Seedance or Kling? The controls below don't apply to you; the framework sets them. They matter the moment you open something like WAN or LTX in ComfyUI.

Concept 01
Latent Space

AI works in a compressed mathematical space with thousands of dimensions, not in pixels. Each dimension encodes a different aspect of what makes an image look the way it does. This map is a tiny 2D slice of something far deeper.

Location
Nearest concept
Hover the map
Blend
Move your cursor to explore. Each cluster represents a learned region, but remember: every point here is actually a vector through thousands of dimensions. Colour, texture, lighting, composition, subject matter, mood: each gets its own axis.
What it is

A compressed map of everything the model has ever learned, held as mathematical coordinates across thousands of axes. No images are stored in it.

The dimensions

One dimension might encode "how warm the lighting is." Another: "how much motion blur." Another: "how close the subject is." Thousands of these overlap and interact.

The cost of the trip

Entering latent space to edit an image always changes it slightly, because you're translating into a different language and back. The more you ask it to change, the more it drifts.

Concept 02
How Generation Works

AI generates by removing noise, guided by your prompt toward a specific location in latent space. Every generated image starts as pure static.

Progress 0%
Pure noise: the starting point
Traditional image creation

Start from blank. Build up pixels using geometry, colour, or brush strokes.

Deterministic. Same input, same result every time.

What's stored is pixels. A colour value at a position. That's the whole file.

Generative AI (diffusion)

Start from pure noise. Every pixel is random, like TV static.

Probabilistic. Same prompt, different result, because it's navigating a space.

What's stored is understanding. The image is decoded out at the end of the process.

Concept 03
The Controls: Steps, CFG, Denoise & ControlNet

These are the dials you get in open-weight tools like WAN and LTX (typically in ComfyUI). On commercial models the framework sets them for you. Understanding what each does means you can stop guessing and start steering. Safe starting points: Steps 25, CFG 7, Denoise 0.5 for edits, ControlNet 0.7. Screenshot those, then read on for why.

3A  ·  Steps
How many rounds of denoising the model runs. More steps = more refined result, up to a point. After around 30–40 steps most models stop improving and just burn compute.
Too few (under 10)

Blurry, unresolved. The denoising hasn't had enough passes to converge on a coherent image. Like developing a photo and pulling it out of the chemical bath too early.

Sweet spot (20–35)

Most models are fully resolved here. Details are sharp, the image is coherent. Going higher usually costs time without visible gain.

Too many (50+)

Diminishing returns. The image has already converged, and extra steps can actually introduce subtle over-sharpening or artefacts in some models.

3B  ·  CFG (Classifier-Free Guidance)
CFG controls how strictly the model sticks to your guidance, whether that's a text prompt or a reference image. Think of it as the dial between "loosely inspired by" and "as literally as possible." Too low and the model ignores you. Too high and it overcooks the image, pushing contrast and colour to extremes.
CFG Scale 7
What it's steering toward
📝
Text prompt: "A lone hilltop at sunset, low sun over distant hills, warm sky fading to dusk, wide landscape"
When using a text prompt, CFG measures how literally the model tries to match your description.
Prompt adherence
Med
Creative freedom
Med
Overcooked risk
Low
Visual result
CFG 7: the default sweet spot for most models. The image follows the prompt clearly without losing natural variation.

CFG also applies to image reference (img2img). When you give the model an image to work from instead of a text prompt, CFG controls how strictly it sticks to the composition, colours, and content of that reference. Low CFG = loosely inspired. High CFG = tries to copy it exactly, then overcooks.

3C  ·  Denoise Strength
Denoise strength controls how much of your input image gets destroyed before generation begins. At 1.0, the model ignores your image completely and generates from scratch. At 0.1, it barely touches it: tiny adjustments only. This is the primary dial for img2img editing.
Denoise 0.50
Input image

Your starting point. Provide your own image; this is a placeholder.

Noise injected

At 0.50, half the image is destroyed

After generation

Result retains structure, varies details

Low (0.1 – 0.3)

Very subtle changes. Good for texture tweaks, style nudges, or colour shifts without changing composition. The image barely moves.

Mid (0.4 – 0.7)

The model reconstructs significant areas while keeping the general layout. Good for lighting changes, subject adjustments, environment swaps.

High (0.8 – 1.0)

The input image is mostly or entirely replaced. At 1.0 it's pure generation: your input image is irrelevant. The prompt takes over completely.

3D  ·  ControlNet
ControlNet gives the model a structural reference such as a depth map, an edge outline or a pose skeleton, and uses it to constrain where things can go. The generation still happens normally, but it has to stay inside the lines. The strength slider controls how strictly those lines are enforced.
CN Strength 0.70
ControlNet input

An edge map or depth map extracted from a reference image. The model uses this as its structural constraint, reading shape and space while ignoring colour and content.

Text prompt

What you want it to look like. Without ControlNet, the model would generate this freely. With ControlNet, it has to fit the structure.

Result

Strength 0.70: structure is clearly enforced, prompt drives the look

What ControlNet is used for

Matching a camera angle from a reference shot. Keeping a character in the same pose across frames. Applying a new style to an existing composition. Making AI results predictable and repeatable.

Strength range

Below 0.3, the structure barely influences the result. Above 0.9, the model is almost tracing the edges and creative output is heavily constrained. 0.6–0.8 is usually where it feels controlled but not robotic.

Concept 04
Why Prompts Work Differently for Generation vs. Editing

Long detailed prompts help when generating from scratch. Short focused prompts are better when editing. The reason comes down to noise, and how much of it you're injecting.

Works well

"A detective at a rain-soaked phone box in 1970s London, night, neon reflections on wet tarmac, wide shot, film grain, cold blue tones"

Every detail steers the model toward a more specific location in latent space. More words = tighter destination.

Too vague

"A man standing in a city at night"

The model has millions of valid locations this could land. You'll get something, but you've surrendered the steering.

When generating, there is no existing image to protect, so more detail is always better.
Concept 05
Why AI Works in 8-Bit

Everything AI generates is standard 8-bit, the same as a JPEG. This limits how far you can push it in a grade. Below is the reason, and the rare exceptions.

8-bit (standard AI)

256 levels per channel. Smooth gradients can show banding. Fine for display, limited for serious grading.

16-bit (LTX LumiVid LoRA)

65,536 levels per channel. Full grading latitude. Matches professional camera output, though it remains rare in AI video.

Why 8-bit?

Training data is almost entirely 8-bit images from the internet. The model learned a world that only has 256 brightness levels per channel.

Why it won't change quickly

16-bit images are rare online. Retraining on a proper high bit-depth dataset requires sourcing an entirely different corpus, which is a massive undertaking.

The LTX Trick

A LoRA teaches the model to output values in log space, so those 256 levels decode into a much wider tonal range, mimicking 16-bit latitude.

Practical impact for filmmakers: Standard AI footage grades like a phone video. Fine for a basic pass, but it falls apart if you push the colour hard. If you're compositing AI elements into real camera footage, this mismatch shows up fast. Until HDR models become standard, treat AI footage like 8-bit MP4: grade gently, and match it to your camera rather than the other way around.

Now apply it
Diagnose the Bad Generation

This is the skill that matters most day to day: looking at a broken result and knowing which dial caused it. Each example shows a real failure mode. Pick the culprit.

Which dial is wrong?
Reference
Symptom → Cause → Fix

The fastest path from "something's wrong" to "here's the dial." Screenshot this one too.

You see
Likely cause
Fix
Fried, over-saturated, harsh contrast
CFG too high
Lower CFG toward 6–8. It's trying too hard to obey the prompt.
Vague, washed out, ignores your prompt
CFG too low
Raise CFG toward 7. The model isn't listening hard enough.
Blurry, soft, unfinished-looking
Steps too low
Raise steps to 20–30. It didn't get enough passes to resolve.
Edit changed the whole image, not just the part you wanted
Denoise too high
Lower denoise toward 0.3–0.5. Or shorten your prompt.
Edit did nothing / barely changed
Denoise too low
Raise denoise toward 0.6. It's protecting too much of the original.
Repeated faces, duplicated objects, warped layout
Resolution above native
Generate at native size, then upscale separately.
Result ignores the reference pose or composition
ControlNet too weak
Raise CN strength toward 0.7–0.8.
Result looks stiff, traced, no creative variation
ControlNet too strong
Lower CN strength toward 0.5–0.6.
Shoot First · The Craft shootfirst.art