Siirry päänavigointiin Siirry hakuun Siirry pääsisältöön

Uimt: A Framework for Improving Unimodal Inference via Multimodal Training

Tutkimustuotos: KonferenssiartikkeliTieteellinenvertaisarvioitu

Abstrakti

The field of multimodal learning is developing rapidly, with emergence of many novel models and applications. Still, works proposing unimodal and multimodal models are generally disjoint, and either focus on fully-unimodal or fully-multimodal scenarios. Nevertheless, oftentimes in real-world applications data of multiple modalities are available during training while only one of them can be utilized during inference due to associated computational costs, or complexity of utilizing additional sensors. In this work, we develop a framework for improving inference of arbitrary unimodal models with multimodal training, without incurring any additional computational cost at inference time, but benefiting from the advantages of multimodal training. We show that our framework is applicable to different architecture types: transformers, 3D CNNs, and 2D+1D CNNs. To showcase this generality we evaluate our approach on tasks of dynamic hand gesture recognition based on RGB and Depth, audiovisual emotion recognition based, and audio-video-text based sentiment analysis. Our approach consistently outperforms the conventionally trained unimodal counterparts. We additionally investigate how within our framework training of multimodal models can benefit from unimodal, modality-specific learning signals. Utilizing the same variety of architectures as mentioned above, we show how models trained with additional supervision from each isolated modality outperform a multimodal-only counterpart.
AlkuperäiskieliEnglanti
Otsikko2024 IEEE International Conference on Image Processing (ICIP)
KustantajaIEEE
Sivut694-700
ISBN (elektroninen)979-8-3503-4939-9
DOI - pysyväislinkit
TilaJulkaistu - 2024
OKM-julkaisutyyppiA4 Artikkeli konferenssijulkaisussa
TapahtumaIEEE International Conference on Image Processing - Abu Dhabi, Yhdistyneet arabiemiirikunnat
Kesto: 27 lokak. 202430 lokak. 2024

Julkaisusarja

NimiProceedings : International Conference on Image Processing
ISSN (elektroninen)2381-8549

Conference

ConferenceIEEE International Conference on Image Processing
Maa/AlueYhdistyneet arabiemiirikunnat
Kaupunki Abu Dhabi
Ajanjakso27/10/2430/10/24

Julkaisufoorumi-taso

  • Jufo-taso 1

Sormenjälki

Sukella tutkimusaiheisiin 'Uimt: A Framework for Improving Unimodal Inference via Multimodal Training'. Ne muodostavat yhdessä ainutlaatuisen sormenjäljen.

Siteeraa tätä