Understanding Generative AI
What it actually does, and why that changes how you use it
Before any control below matters, know which kind of model you're using. It decides whether you'll ever touch these dials.
The service wraps the model and runs it for you. You prompt; it handles the rest.
- Prompt tuning, enhancement, and upscaling happen automatically
- Reliable results with almost no setup
- The controls in this sheet are handled behind the scenes
You run the model yourself, usually in ComfyUI, with every dial exposed.
- Where Steps, CFG, Denoise, and ControlNet live
- Where LoRAs, custom pipelines, and deep tinkering happen
- Higher ceiling once you put the hours in
Commercial wins on reliability out of the box. Open-weight wins on control and ceiling once you invest the time. Pick the one that matches how you want to work.
On Seedance or Kling? The controls below don't apply to you; the framework sets them. They matter the moment you open something like WAN or LTX in ComfyUI.
AI works in a compressed mathematical space with thousands of dimensions, not in pixels. Each dimension encodes a different aspect of what makes an image look the way it does. This map is a tiny 2D slice of something far deeper.
A compressed map of everything the model has ever learned, held as mathematical coordinates across thousands of axes. No images are stored in it.
One dimension might encode "how warm the lighting is." Another: "how much motion blur." Another: "how close the subject is." Thousands of these overlap and interact.
Entering latent space to edit an image always changes it slightly, because you're translating into a different language and back. The more you ask it to change, the more it drifts.
AI generates by removing noise, guided by your prompt toward a specific location in latent space. Every generated image starts as pure static.
Start from blank. Build up pixels using geometry, colour, or brush strokes.
Deterministic. Same input, same result every time.
What's stored is pixels. A colour value at a position. That's the whole file.
Start from pure noise. Every pixel is random, like TV static.
Probabilistic. Same prompt, different result, because it's navigating a space.
What's stored is understanding. The image is decoded out at the end of the process.
These are the dials you get in open-weight tools like WAN and LTX (typically in ComfyUI). On commercial models the framework sets them for you. Understanding what each does means you can stop guessing and start steering. Safe starting points: Steps 25, CFG 7, Denoise 0.5 for edits, ControlNet 0.7. Screenshot those, then read on for why.
Blurry, unresolved. The denoising hasn't had enough passes to converge on a coherent image. Like developing a photo and pulling it out of the chemical bath too early.
Most models are fully resolved here. Details are sharp, the image is coherent. Going higher usually costs time without visible gain.
Diminishing returns. The image has already converged, and extra steps can actually introduce subtle over-sharpening or artefacts in some models.
CFG also applies to image reference (img2img). When you give the model an image to work from instead of a text prompt, CFG controls how strictly it sticks to the composition, colours, and content of that reference. Low CFG = loosely inspired. High CFG = tries to copy it exactly, then overcooks.
Your starting point. Provide your own image; this is a placeholder.
At 0.50, half the image is destroyed
Result retains structure, varies details
Very subtle changes. Good for texture tweaks, style nudges, or colour shifts without changing composition. The image barely moves.
The model reconstructs significant areas while keeping the general layout. Good for lighting changes, subject adjustments, environment swaps.
The input image is mostly or entirely replaced. At 1.0 it's pure generation: your input image is irrelevant. The prompt takes over completely.
An edge map or depth map extracted from a reference image. The model uses this as its structural constraint, reading shape and space while ignoring colour and content.
What you want it to look like. Without ControlNet, the model would generate this freely. With ControlNet, it has to fit the structure.
Strength 0.70: structure is clearly enforced, prompt drives the look
Matching a camera angle from a reference shot. Keeping a character in the same pose across frames. Applying a new style to an existing composition. Making AI results predictable and repeatable.
Below 0.3, the structure barely influences the result. Above 0.9, the model is almost tracing the edges and creative output is heavily constrained. 0.6–0.8 is usually where it feels controlled but not robotic.
Long detailed prompts help when generating from scratch. Short focused prompts are better when editing. The reason comes down to noise, and how much of it you're injecting.
"A detective at a rain-soaked phone box in 1970s London, night, neon reflections on wet tarmac, wide shot, film grain, cold blue tones"
Every detail steers the model toward a more specific location in latent space. More words = tighter destination.
"A man standing in a city at night"
The model has millions of valid locations this could land. You'll get something, but you've surrendered the steering.
Everything AI generates is standard 8-bit, the same as a JPEG. This limits how far you can push it in a grade. Below is the reason, and the rare exceptions.
256 levels per channel. Smooth gradients can show banding. Fine for display, limited for serious grading.
65,536 levels per channel. Full grading latitude. Matches professional camera output, though it remains rare in AI video.
Training data is almost entirely 8-bit images from the internet. The model learned a world that only has 256 brightness levels per channel.
16-bit images are rare online. Retraining on a proper high bit-depth dataset requires sourcing an entirely different corpus, which is a massive undertaking.
A LoRA teaches the model to output values in log space, so those 256 levels decode into a much wider tonal range, mimicking 16-bit latitude.
Practical impact for filmmakers: Standard AI footage grades like a phone video. Fine for a basic pass, but it falls apart if you push the colour hard. If you're compositing AI elements into real camera footage, this mismatch shows up fast. Until HDR models become standard, treat AI footage like 8-bit MP4: grade gently, and match it to your camera rather than the other way around.
This is the skill that matters most day to day: looking at a broken result and knowing which dial caused it. Each example shows a real failure mode. Pick the culprit.
The fastest path from "something's wrong" to "here's the dial." Screenshot this one too.
