VOID is an open-source AI model developed by Netflix designed for advanced video object removal that accounts for physical interactions, such as items falling after an object is deleted. It utilizes a CogVideoX 3B Transformer architecture with specialized quadmask conditioning to maintain scene consistency.
Highlights
Removes objects along with their induced physical effects like shadows, reflections, and displaced items.
Uses a 'quadmask' system (remove, overlap, affected, background) for precise spatial conditioning.
Features a two-pass inference process where the second pass uses optical flow-warped latent initialization to improve temporal consistency.
Requires high-end hardware with at least 40GB of VRAM (e.g., NVIDIA A100).