Survey of Research on Multimodal Fusion Technology for Deep Learning
摘要
Multimodal Fusion Technology(MFT) for Deep Learning(DL) refers to the conversion and fusion of information obtained by machine from texts,images,voices,videos and other materials,so as to improve the performance of the model.The universality of modals and the heat of DL boost the rapid development of multimodal fusion.In order to improve the performance of DL model classification or regression,this paper summarizes the multimodal fusion architecture,fusion methods and alignment technologies in the early stage of MFT development.This paper focuses on the analysis of the three fusion architectures:joint,cooperative and codec architectures,in terms of their adoption in DL and advantages/disadvantages.The specific fusion methods and alignment technologies such as Multiple Kernel Learning(MKL),Graphic Model(GM) and Neural Network(NN) are also studied.Finally,the public datasets commonly used in multimodal fusion research are summarized,and the direction of further research in cross-modal transfer learning,resolution of modal semantic conflicts,and multimodal combination evaluation is prospected.