In the proposed project, approaches will be implemented towards the detection of emotions by fusing the relevant details from audio, video, and audio-visual modalities. As part of the proposed work, state-of-the-art automatic emotion detection tools/techniques will be studied, whereby some peculiar artifacts and gaps uncovered in the literature will be analyzed and acted upon. The proposed technique would exploit multimodal and multilingual information, trained with Deep Learning architectures to discriminate various emotions and behaviors with greater accuracy.