Data/ML/AIEmergingAround since 2023
Multimodal AI
AI models that can understand and generate across multiple types of input/output — text, images, audio, video — rather than just one.
Think of it like
Like a person who can look at a photo, read the caption, and listen to the audio all at once and make sense of them together.
Junior or senior?
Junior sounds like
Uses separate text and image tools side by side and calls it multimodal.
Senior sounds like
Can describe a feature that actually combined more than one input type together.
Ask them
“What's an example of a feature you built that combined more than one type of input, like text and images?”