Multimodal models are AI systems trained to take in and/or produce more than one kind of data — text, images, audio, video — in a single model, rather than requiring separate models stitched together.
Continue to AI University →