Multimodal Model Limits: What Text-Image-Video Understanding Can and Can’t DoImage-Text Understanding, Multimodal AI, Video Understanding