01

More than one modality

A modality is a form of information, such as text, images, audio, video, sensor readings, or actions. A multimodal system accepts, relates, or produces more than one of these forms.

Examples include answering a question about a photograph, transcribing and summarizing speech, generating an image from text, or relating video frames to written instructions.

02

How information is connected

Models convert each modality into numerical representations. Training aligns or combines those representations so related content can influence a shared task. Some systems use separate encoders connected to a language model; others integrate modalities more deeply.

The exact architecture matters. Two products described as multimodal may have very different capabilities, context limits, and failure modes.

03

What multimodality enables

Combining modalities can improve accessibility, document understanding, robotics, scientific analysis, creative tools, and interfaces that better match how people communicate.

It also creates new privacy and security concerns because images, voices, locations, faces, screens, and surroundings may reveal sensitive information.

04

Limits across modalities

A model may read text well but misinterpret a chart, recognize objects but miss spatial relationships, or confidently describe something absent from an image. Success in one modality does not transfer automatically to another.

A category describes how a system is built or used; it does not by itself prove accuracy, safety, intelligence, or consciousness. Real performance must be tested on the actual task and conditions of use.

SOURCES

Primary research and institutions

Sources are linked directly so readers can examine the underlying evidence. Numerical statements identify the reporting organization and year. Projections are presented as estimates, not established future facts.