DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
arXiv, 2026

Flow-matching video generators produce temporally coherent output yet routinely violate elementary physics, because reconstruction objectives penalise per-frame deviation without distinguishing consistent dynamics from impossible ones. DiReCT identifies semantic-physics entanglement as the obstacle to contrastive flow matching in text-conditioned video, then splits the contrastive signal into a macro term that draws negatives from semantically distant regions and a micro term whose hard negatives share full scene semantics but differ along a single axis of physical behaviour.
