What the Genjutsu model is
Genjutsu is a reference-driven video-to-video model built for two jobs: transferring the motion of an existing clip onto new references, and swapping individual elements inside a shot. It is not a text-to-video model, and that distinction explains most of how it behaves.
Because the source clip supplies the motion, the camera and the timing, the model can spend its capacity on appearance: who and what is in the frame, and how they are lit. That is why results stay stable across a whole shot instead of drifting the way prompt-only video generators often do.