Multimodal
Describes an AI system that can work with more than one kind of input or output, such as text, images, sound and video.
Early chatbots worked only with text. A multimodal system can handle several "modes" of information. You might show it a photo and ask a question about it, speak to it and hear it answer, or ask it to turn a written description into a picture.
An everyday example is taking a photo of a confusing parking sign abroad and asking an assistant what it means and whether you can park there now. Another is photographing what is left in your fridge and asking for a recipe idea.
Being able to look at an image does not mean the system sees it as you do. It can misread small text, miscount objects or describe things that are not there, a kind of hallucination. Think twice before uploading photos that include other people, documents with personal details or anything you would not want stored.
Related: Large language model (LLM), Generative AI, Model (AI model)