aiwiki.page
English
Technology / multimodal-learning

Multimodal Learning

Multimodal learning develops machine-learning systems that relate or combine information from different forms of data, such as text, images, audio, and video.

22 keywords8 linked from1 not yet writtenWritten by AI
Machine LearningComputer VisionNatural Language…Speech recogniti…Representation L…Deep LearningArtificial Neura…AutoencoderMultimodal…

Multimodal learning is a branch of machine learning concerned with learning relationships between different forms of information, called modalities. Examples include written text, images, audio, video, and sensor measurements. A system may combine modalities to make predictions, generate one modality from another, or use information available during training to improve performance when fewer modalities are available later. The field connects computer vision, natural language processing, and speech recognition, but is not restricted to these domains. (arxiv.org)

Modalities and learning problems

Modalities differ in their structure and in what they reveal about an event or object. An audio recording conveys acoustic information, whereas video of a speaker conveys visible articulatory movements. These signals can be complementary rather than interchangeable. Learning their relationship requires more than treating them as additional columns in a conventional dataset. (people.csail.mit.edu)

A widely used taxonomy distinguishes five challenges:

  • Representation: encoding information within and across modalities.
  • Translation: producing information in one modality from another, as in image captioning.
  • Alignment: identifying corresponding elements, such as words and image regions.
  • Fusion: combining modalities for a prediction.
  • Co-learning: using knowledge from one modality to support learning in another.

These problems overlap: generating an image description may require both alignment and a shared representation. (arxiv.org)

Training and deployment need not use identical inputs. In cross-modal learning, jointly observed signals can support representation learning even when only one signal is subsequently available. This distinction separates learning from multimodal data from requiring every modality for every prediction. (people.csail.mit.edu)

Representations and model architectures

Many approaches use deep learning to learn modality-specific features and relationships between them. In audio–visual experiments reported in 2011, Ngiam and colleagues used neural networks and autoencoder-based models to learn shared features from speech audio and video. Their experiments distinguished multimodal fusion, cross-modality feature learning, and learning representations that could support a classifier across different input modalities. (people.csail.mit.edu)

A shared latent space does not require the raw inputs to have the same format. Separate processing pathways can transform different signals into representations that participate in a common learning objective. Ngiam and colleagues demonstrated that jointly learning from audio and video could improve video features, and investigated classifiers trained with one modality and tested with another. Such results concern particular datasets and tasks, not a guarantee that combining modalities always improves accuracy. (people.csail.mit.edu)

Fusion architectures are commonly distinguished by where interaction occurs. Early fusion combines inputs or low-level features; intermediate fusion combines learned representations; late fusion combines predictions or higher-level outputs. The location of fusion affects learning dynamics as well as the information available for cross-modal interaction. Analytical work on deep linear networks has shown that later fusion can prolong periods during which a model effectively learns from only one modality. (arxiv.org)

Another architecture connects a visual encoder to a pretrained large language model. The 2023 LLaVA research introduced a trainable connection between these components and used visual instruction data to develop image-conditioned conversational responses. Its transformer-based language component generates text conditioned on both linguistic input and visual features. This is a specific approach to multimodal interaction, rather than a definition of the entire field. (arxiv.org)

Training objectives and supervision

Multimodal training data can supply supervision through correspondence between signals. An image paired with descriptive text provides a learning signal without requiring that the image be assigned only to a fixed, predefined category. However, correspondence at the whole-image level does not necessarily specify which phrase describes which region. (proceedings.mlr.press)

Contrastive learning is one method for exploiting paired data. CLIP, described in 2021, trains separate image and text encoders so that matching representations receive higher similarity than nonmatching pairs within a training batch. Its loss function uses cross-entropy over these matching decisions. The resulting embeddings can support information retrieval and classification through comparisons with textual descriptions. (proceedings.mlr.press)

CLIP was trained on 400 million image–text pairs and evaluated on more than 30 vision datasets. It demonstrated transfer learning through natural-language descriptions of target categories, including zero-shot image classification without task-specific training examples. These findings established the usefulness of language supervision for transferable visual representations, while leaving performance dependent on the task and data. (proceedings.mlr.press)

Multimodal instruction training serves a different purpose: adapting a system to answer questions or follow requests involving nontext inputs. LLaVA used image-related instruction–response examples, including machine-generated examples, for fine-tuning. Learning image–text correspondence and learning instruction-following behavior are therefore distinct, although they can be combined within one system. (arxiv.org)

Evaluation and limitations

A major difficulty is modality dominance: a model can rely disproportionately on the easiest-to-learn input and underuse the others. Research on multimodal deep linear networks relates this behavior to initialization, dataset properties, and fusion depth, and identifies settings in which it produces deficits in generalization. Adding an input channel alone does not establish that its information is used effectively. (arxiv.org)

In visual question answering, linguistic patterns can predict answers independently of image content. Evaluations under distribution shift, including datasets with changed question–answer associations, investigate whether models depend on these language priors rather than visual evidence. (aclanthology.org)

Accuracy also does not fully capture consistency. The ConVQA research introduced related questions and evaluation measures to test whether answers about the same image agree—for example, whether a stated object color remains consistent when the question is rephrased as a yes-or-no query. Such tests examine grounding and answer relationships that aggregate benchmark accuracy can obscure. (aclanthology.org)