Self-Supervised Joint Discovery of Fusion Taxonomies and Multimodal Fusion Strategies for Sentiment Analysis

Main Article Content

Bharathi Niruti, Sujith AVLN, Kanaka Durga Returi

Abstract

Prior studies on multimodal fusion have built two distinct aspects: taxonomy-aware systems which identify the type of fusion strategy (early, late, hybrid, attention-based or graph-based), and fusion architectures which perform multimodal fusion for a specific downstream task and are not dependent on the taxonomy. The Taxonomy-Guided Fusion Strategy Selector (TG-FSS) fuses these streams, leveraging an externally supervised taxonomy signal to boost the accuracy of sentiment classification, F1, and correlation with human judgments over the best prior systems on CMU-MOSEI, CMU-MOSI, and the Multimodal Fusion Strategy Corpus (MFSC), and it uses segment-level and taxonomy annotations to train its taxonomy-alignment component. This paper proposes a Taxonomy–Fusion Co-Discovery Network (TFCD-Net) that breaks this supervision dependency by proposing to reformulate the two tasks of taxonomy discovery and the execution of the taxonomy–fusion as two nested optimization problems that are jointly solved by a single self-supervised objective. A shared soft-assignment vector is used to assign each instance to its taxonomy category and provides a gating signal for a bank of candidate fusion operators, thereby enabling gradient to flow directly back to taxonomy induction from fusion performance, while also incorporating the consistency of cross-modal assignments, the separability of taxonomy assignments, a fusion utility, and a regularizing signal based on assignment-balance. Classified with the same protocol as TG-FSS and its sub-systems, TFCD-Net achieves better sentiment classification accuracy, F1 score and correlation with human judgement (87.1% vs. 85.4%, 0.85 vs. 0.83, 0.84 vs. 0.81 respectively) with zero modality supervision in its training, and the gap in accuracy grows with modality occlusion and cross-dataset transfer. Ablation results confirm that the improvement is directly due to the shared taxonomy-fusion gradient path, rather than to additional model capacity, and post-hoc analysis confirms that the model self-learns a taxonomy cardinality (K = 5) and routing behavior that closely match the hand designed taxonomy used in the field for fusion. The results show that taxonomy discovery and execution of the taxonomy-guided pipelines are two distinct problems, but when they are optimized together, they are reinforcing each other, and that reinforcement can be achieved without annotation cost which is needed for supervised taxonomy-guided pipelines.

Article Details

Section
Articles