Back to the index
35/ 55

FOUNDATIONS

Multimodal.

Multimodal AI · modality

Able to work with more than one kind of information, such as text, images, or audio.

In plain words

A modality is a type of information or signal. A multimodal model or system works across two or more modalities, potentially connecting information in text, images, audio, video, or other data.

A closer look

Different systems support different combinations of input and output. A model may accept an image and text but respond only in text. Another may generate speech. Video can involve both a sequence of images and sound. A product’s multimodal capabilities may also come from several connected models.

The interesting capability is using one kind of information to interpret another: answering a question about a chart, describing an image, or relating spoken language to a visual scene. These tasks introduce their own failure modes, such as reading small text incorrectly or missing the timing of an event.

In practice

AN EXAMPLE

You upload a photograph of a handwritten shopping list and ask for it to be organized by grocery aisle. The system combines visual recognition with language processing.

A useful distinction

Multimodal does not mean equally capable in every modality. Being able to accept an image does not imply precise spatial reasoning, and receiving video does not guarantee that every frame is processed or remembered.

Sources & further reading

Baltrušaitis et al. — Multimodal Machine Learning: A Survey and Taxonomy (opens in a new tab)