Skip to main navigation Skip to search Skip to main content

Uimt: A Framework for Improving Unimodal Inference via Multimodal Training

Research output: Chapter in Book/Report/Conference proceedingConference contributionScientificpeer-review

Abstract

The field of multimodal learning is developing rapidly, with emergence of many novel models and applications. Still, works proposing unimodal and multimodal models are generally disjoint, and either focus on fully-unimodal or fully-multimodal scenarios. Nevertheless, oftentimes in real-world applications data of multiple modalities are available during training while only one of them can be utilized during inference due to associated computational costs, or complexity of utilizing additional sensors. In this work, we develop a framework for improving inference of arbitrary unimodal models with multimodal training, without incurring any additional computational cost at inference time, but benefiting from the advantages of multimodal training. We show that our framework is applicable to different architecture types: transformers, 3D CNNs, and 2D+1D CNNs. To showcase this generality we evaluate our approach on tasks of dynamic hand gesture recognition based on RGB and Depth, audiovisual emotion recognition based, and audio-video-text based sentiment analysis. Our approach consistently outperforms the conventionally trained unimodal counterparts. We additionally investigate how within our framework training of multimodal models can benefit from unimodal, modality-specific learning signals. Utilizing the same variety of architectures as mentioned above, we show how models trained with additional supervision from each isolated modality outperform a multimodal-only counterpart.
Original languageEnglish
Title of host publication2024 IEEE International Conference on Image Processing (ICIP)
PublisherIEEE
Pages694-700
ISBN (Electronic)979-8-3503-4939-9
DOIs
Publication statusPublished - 2024
Publication typeA4 Article in conference proceedings
EventIEEE International Conference on Image Processing - Abu Dhabi, United Arab Emirates
Duration: 27 Oct 202430 Oct 2024

Publication series

NameProceedings : International Conference on Image Processing
ISSN (Electronic)2381-8549

Conference

ConferenceIEEE International Conference on Image Processing
Country/TerritoryUnited Arab Emirates
City Abu Dhabi
Period27/10/2430/10/24

Publication forum classification

  • Publication forum level 1

Fingerprint

Dive into the research topics of 'Uimt: A Framework for Improving Unimodal Inference via Multimodal Training'. Together they form a unique fingerprint.

Cite this