R&D / AI VFX · Relighting

Building a Defensible AI Relighting Pipeline

A waveform-first production SOP for separating physically motivated light from hallucinated exposure, unstable shadows and mathematical artifacts.

Pipeline concept validated · Model tuning in progress

Before/after AI relighting of a Toronto interior — hero placeholder while the final comparison video is in production.

Hero slot: a seamless before/after relighting comparison. The final media is a muted, looping, inline video with a poster and a reduced-motion fallback; until it is delivered, this branded frame stands in so nothing reads as broken.

Relighting a shot after the fact, lifting a flat interior into warm key-and-fill or turning midday into magic hour, is one of the highest-leverage things generative tools can do for a production. It can save a reshoot. It also introduces a trust problem: a model can produce something that looks relit while quietly inventing brightness no real light could create. On a single hero frame you might get away with it. Across a moving sequence, the eye eventually catches the lie.

So at Cinematechs we stopped grading AI relights by eye and started reading them on a waveform. The question is not "does this look nice?" It is "could light have done this to the signal?" This report lays out the standard operating procedure we use to answer that, and it is honest about where the method is proven and where it is still being tuned.

Executive verdict

Pipeline concept Valid
Evaluation method Valid · waveform-first
Current output Optimization required
Production position Advanced R&D · standardizing

Read that carefully, because the two things people conflate are separate. The method for evaluating relighting is sound and repeatable. The model output still needs work: highlights that occasionally clamp, shadows that drift, sequences that breathe. Those are model-control and stabilization problems, not evidence that the underlying pipeline concept is wrong. We are treating this as a validated methodology still undergoing numerical tuning, and we are not going to publish "market leader" as if it were an independently verified fact. It is not one yet. The method is.

Establishing the ground truth

Every test starts from a controlled baseline, because you cannot judge a relight against a memory. The chain is deliberately boring:

  1. A known Rec.709 daylight source. A clean plate with documented exposure and white balance, so the starting signal is not in dispute.
  2. Intentional relighting logic. We ask for a specific, motivated change, a warmer key from frame left, a lower fill, not "make it cinematic."
  3. Temporal video generation. The relight is produced across a moving sequence, not a single frame, because motion is where instability hides.
  4. High-resolution mezzanine output. The result is written to a high-bit-depth container for finishing, kept separate from any web derivative.
  5. Scope-based comparison against the original truth plate, on a waveform and parade, not by eye.

One nuance matters here. When the output differs from the plate, we do not automatically call that a grading error. A different result may simply be a different lighting environment, which is the whole point of relighting. The job is to decide whether the difference is one light could produce, or one only mathematics could.

Contact sheet placeholder: source plate, relight pass, and mezzanine output side by side — final media in production.
Pipeline contact sheet (in production): the source plate, the relight pass, and the mezzanine master read left to right, each with its own scope. Illustrative placeholder pending the final extraction.

The physics lens

Here is the governing rule, and it is the sentence the entire SOP hangs from:

Different lighting is allowed to change the signal. It is not allowed to break it.

Real illumination redistributes energy in lawful ways. It lifts midtones, opens or deepens shadow, separates warm and cool sources, bends the contrast curve. What it never does is manufacture a single-pixel-wide spike of brightness out of nothing. When you see that on a scope, you are not looking at a highlight. You are looking at a calculation error. The interactive below toggles between the two cases so the difference is legible rather than asserted.

Interactive · illustrative
Toggle the two cases. "Physical light" shows a motivated midtone lift with a gently thickening highlight band, energy spreading the way light does. "Math artifact" shows the tell: a thin needle spike with no width. Traces are illustrative, keyboard operable, and static under reduced-motion.

What light is allowed to change

Four behaviours are always legitimate, because each one moves a population of pixels together rather than inventing isolated structure:

  • Midtone movement from an exposure change, the whole tonal mass sliding up or down.
  • Shadow lifting or lowering from a new key-to-fill ratio, with the black floor still anchored where it belongs.
  • RGB separation from mixed colour temperatures, a warm key reading against a cool ambient, which is how real multi-source scenes behave.
  • Contrast-curve modification from a new motivated source, the tonal slope bending as a different light takes over.

On a scope these show up as bands shifting, gradients tilting and highlights rolling off. They do not show up as isolated, structureless spikes. If the change has population and roll-off, it is probably light. If it is a lonely vertical line, it is probably not.

The highlight test

If you only remember one diagnostic from this report, remember this one, because it catches more fakes than any other single check.

Highlights must expand, not spike. A thin line is a calculation error; a thick band behaves like a light source.

  • Pass: the highlight expands into a thick energy band with a bright core and a visible roll-off, because real light falls off.
  • Fail: the highlight appears as a thin needle, a hard plateau, or an isolated spike with no roll-off, a delta function the model painted in.

Genuine specular energy has width. An invented highlight does not. Catch the needle and you have caught the hallucination, before it reaches a client review rather than after. This is the same interactive as above, examined specifically for highlight shape.

Shadows and time

A locked camera pointed at a static lighting environment should produce a stable waveform. If nothing in the scene is moving and nothing about the light is changing, the trace should sit still. When it does not, the model is telling on itself. The failure modes we watch for:

  • Shadow-channel crossover, where the RGB channels cross in the toe and neutral blacks pick up a colour cast.
  • Unmotivated RGB drift with no source to justify it.
  • Black-floor movement, the darkest values sliding frame to frame.
  • Pulsing exposure, the overall level breathing in and out.
  • Frame-to-frame highlight changes with no moving specular source.
  • Temporal "breathing," a slow instability that a single frame will never reveal.

The interactive below holds the scene static and lets you compare a stable scope against an unstable one. Watch the band: stable stays put; unstable pulses and drags its black floor.

Interactive · illustrative
Static scene, two outcomes. "Stable" holds a steady band. "Unstable" shows temporal breathing and a drifting black floor (red line), the signature of an AI stability failure. Motion pauses to a representative static frame under reduced-motion.
Shadow-channel crossover placeholder: RGB parade showing channels crossing in the toe — final media in production.
Shadow-channel crossover (in production): an RGB parade where the channels cross in the toe, turning a neutral black slightly warm or cool. A subtle, common, and disqualifying failure. Illustrative placeholder pending capture.

Pass / borderline / fail matrix

This is the working rubric a supervisor uses on a review. Status is carried by an explicit label and an icon, never by colour alone.

Pass
  • Motivated exposure differences
  • Plausible colour-temperature (CCT) shifts
  • Tungsten, daylight or moonlight variants
  • Stable shadow structure over time
  • High-resolution mezzanine delivery
Borderline · tuning
  • Small specular hallucinations
  • Shadow chroma noise
  • Minor temporal breathing
  • Slightly aggressive highlight roll-off
Fail
  • Unmotivated waveform drift
  • Highlight clipping or flat ceilings
  • Thin highlight needles
  • RGB shadow crossover
  • Unstable light when the scene should be static

The Cinematechs relighting SOP

Seven rules govern whether a relight ships. They are the operational form of the physics lens.

  1. Lighting context

    Exposure, contrast and colour temperature may change only when motivated by the intended source. Change without motivation is a red flag, not a style.

  2. Signal safety

    No illegal clipping, no flat highlight ceilings, no shadow-channel crossover. The signal has to stay legal on the master.

  3. Highlight quality

    Highlights expand and roll off. They do not spike or clamp. Width is the evidence that a light source, not a solver, produced them.

  4. Shadow integrity

    Shadows may become darker or lighter, but they must stay chromatically stable. A neutral black should remain neutral unless a coloured source explains otherwise.

  5. Temporal coherence

    Static light requires a stable waveform. Unmotivated breathing is an AI-instability failure, full stop, no matter how good the single frame looks.

  6. Bit-depth truth

    A 4K 16-bit file can be a mezzanine container without holding true 16-bit source precision. We never claim recovered precision that was never captured. If the source began as 8-bit, we use appropriate dithering or grain during finishing rather than pretending the extra bits are real.

  7. Labeling standards

    Outputs are labelled by observed behaviour, not by hope:

    AI gradeglobal exposure or colour remapping.
    AI relightlighting changes that behave physically on the scope.
    Experimentalresults that still contain unresolved temporal or signal failures.

Why the pipeline is not yet "perfect"

Transparency is part of the standard, so here is the honest state of things. The remaining problems are real engineering limitations, not marketing caveats:

  • Models occasionally break physical logic and need to be reined back in.
  • Video generation can introduce temporal instability that a still never shows.
  • Highlight mathematics sometimes replaces believable light behaviour with a solver's guess.
  • Numeric thresholds, the exact IRE tolerances that turn a "borderline" into a "fail," still need to be formalized.
  • More measured sequences are required before we publish any certification claim.

None of that invalidates the method. It is the difference between a validated way of judging the work and a fully industrialised way of guaranteeing it. We are between those two points, and we would rather say so.

Next steps

  1. Define numeric waveform thresholds so pass and fail stop being a judgement call.
  2. Establish model-specific tuning ranges for the generators we actually use.
  3. Lock temporal-stabilization procedures for sequences, not just frames.
  4. Build a repeatable QC checklist a supervisor can run the same way every time.
  5. Test the method across varied footage, interiors, exteriors, mixed sources, skin tones.
  6. Convert the internal SOP into a formal white paper or certification specification once the numbers are locked.

Key takeaways

  • Different lighting may change the signal, but it may not break it. That single rule drives the whole SOP.
  • Judge relights on a waveform and across motion, not on one polished frame.
  • The most reliable fail signature is highlight shape: a thick band passes, a thin needle fails.
  • The method is validated; the model output is still being tuned. We label work by what it actually does.

Does your relight survive the timeline?

Cinematechs develops AI-assisted lighting and VFX workflows that are evaluated in motion and validated on the signal, not judged from a single polished frame.

Discuss a shot