Direct complex multimodal tasks in natural language
Combine up to 9 images, 3 videos, and 3 audio clips within 12 mixed files, then explain their roles in one instruction—for example, reference a camera move, a character image, and a voice or song together.

