TY - GEN
T1 - Uimt
T2 - IEEE International Conference on Image Processing
AU - Chumachenko, Kateryna
AU - Iosifidis, Alexandros
AU - Gabbouj, Moncef
PY - 2024
Y1 - 2024
N2 - The field of multimodal learning is developing rapidly, with emergence of many novel models and applications. Still, works proposing unimodal and multimodal models are generally disjoint, and either focus on fully-unimodal or fully-multimodal scenarios. Nevertheless, oftentimes in real-world applications data of multiple modalities are available during training while only one of them can be utilized during inference due to associated computational costs, or complexity of utilizing additional sensors. In this work, we develop a framework for improving inference of arbitrary unimodal models with multimodal training, without incurring any additional computational cost at inference time, but benefiting from the advantages of multimodal training. We show that our framework is applicable to different architecture types: transformers, 3D CNNs, and 2D+1D CNNs. To showcase this generality we evaluate our approach on tasks of dynamic hand gesture recognition based on RGB and Depth, audiovisual emotion recognition based, and audio-video-text based sentiment analysis. Our approach consistently outperforms the conventionally trained unimodal counterparts. We additionally investigate how within our framework training of multimodal models can benefit from unimodal, modality-specific learning signals. Utilizing the same variety of architectures as mentioned above, we show how models trained with additional supervision from each isolated modality outperform a multimodal-only counterpart.
AB - The field of multimodal learning is developing rapidly, with emergence of many novel models and applications. Still, works proposing unimodal and multimodal models are generally disjoint, and either focus on fully-unimodal or fully-multimodal scenarios. Nevertheless, oftentimes in real-world applications data of multiple modalities are available during training while only one of them can be utilized during inference due to associated computational costs, or complexity of utilizing additional sensors. In this work, we develop a framework for improving inference of arbitrary unimodal models with multimodal training, without incurring any additional computational cost at inference time, but benefiting from the advantages of multimodal training. We show that our framework is applicable to different architecture types: transformers, 3D CNNs, and 2D+1D CNNs. To showcase this generality we evaluate our approach on tasks of dynamic hand gesture recognition based on RGB and Depth, audiovisual emotion recognition based, and audio-video-text based sentiment analysis. Our approach consistently outperforms the conventionally trained unimodal counterparts. We additionally investigate how within our framework training of multimodal models can benefit from unimodal, modality-specific learning signals. Utilizing the same variety of architectures as mentioned above, we show how models trained with additional supervision from each isolated modality outperform a multimodal-only counterpart.
U2 - 10.1109/ICIP51287.2024.10647735
DO - 10.1109/ICIP51287.2024.10647735
M3 - Conference contribution
T3 - Proceedings : International Conference on Image Processing
SP - 694
EP - 700
BT - 2024 IEEE International Conference on Image Processing (ICIP)
PB - IEEE
Y2 - 27 October 2024 through 30 October 2024
ER -