This project investigates the "text predominance" phenomenon in multimodal emotion analysis, where existing models rely excessively on textual information while underutilising audio and visual cues. By combining linguistic insights with interpretable AI, the project aims to quantify the contributions of different modalities, diagnose how multimodal models integrate them, and develop transparent, context-aware fusion methods that more faithfully reflect human emotional communication.