In plain words
A modality is a type of information or signal. A multimodal model or system works across two or more modalities, potentially connecting information in text, images, audio, video, or other data.
A closer look
Different systems support different combinations of input and output. A model may accept an image and text but respond only in text. Another may generate speech. Video can involve both a sequence of images and sound. A product’s multimodal capabilities may also come from several connected models.
The interesting capability is using one kind of information to interpret another: answering a question about a chart, describing an image, or relating spoken language to a visual scene. These tasks introduce their own failure modes, such as reading small text incorrectly or missing the timing of an event.
In practice
You upload a photograph of a handwritten shopping list and ask for it to be organized by grocery aisle. The system combines visual recognition with language processing.
A useful distinction
Multimodal does not mean equally capable in every modality. Being able to accept an image does not imply precise spatial reasoning, and receiving video does not guarantee that every frame is processed or remembered.