Interspeech 2026
1 Plaksha University, India
2 IIIT Delhi, India
3 NTHU, Taiwan
4 University of Tartu, Estonia
In this work, we investigate the robustness of audio–visual (AV) cues for Multilingual Audio–Visual Hate Speech Detection (M-AVHSD). We propose VINAYAKA, a novel framework for M-AVHSD that relies solely on AV cues for decision-making. Along with WavLM and ImageBind, it employs Cross-modal Fusion in Hyperbolic Space (CFHS), which models the effective alignment between paralinguistic and behavioral AV cues to better capture hateful intent in temporal audio signals and corresponding visual scenes. Empirical results demonstrate that the proposed framework achieves state-of-the-art performance in both in-distribution and out-of-distribution settings across cross-lingual and cross-dataset evaluations. Furthermore, the results show that the framework remains robust to ASR error propagation.
Detects hateful intent using only audio-visual cues without relying on transcriptions.
Introduces Cross-modal Fusion in Hyperbolic Space (CFHS) to model interactions between audio and visual behavioral signals.
Generalizes across languages and datasets under challenging out-of-distribution conditions.
Remains robust even when automatic speech recognition systems introduce errors.
VINAYAKA (Multilingual Audio-Visual Hate Speech Detection via Cross-Modal Fusion in Hyperbolic Space) detects hateful intent directly from audio and visual signals without relying on textual transcriptions. The framework combines pretrained audio and visual representations with hyperbolic cross-modal reasoning to model interactions between paralinguistic speech cues and behavioral visual cues. By performing multimodal fusion in hyperbolic space, VINAYAKA captures hierarchical relationships between modalities and improves generalization across languages and datasets.
Raw audio waveforms are processed using WavLM, while video frames are encoded using the ImageBind Visual Encoder. These pretrained encoders generate rich semantic and behavioral representations for each modality. The resulting audio embedding E(a) and visual embedding E(v) capture complementary information relevant to multilingual hate speech detection.
Audio and visual representations are independently passed through a lightweight 1D-CNN followed by max-pooling. This stage refines modality-specific patterns, suppresses noise, and aggregates temporal information. The pooled representations are flattened to produce compact feature vectors suitable for multimodal interaction.
The flattened Euclidean representations are projected into a shared hyperbolic space. Unlike Euclidean geometry, hyperbolic geometry naturally models hierarchical and complex relational structures. The projected embeddings form hyperbolic audio sequences and hyperbolic video sequences that provide a richer representation space for cross-modal reasoning.
VINAYAKA performs bidirectional cross-modal reasoning through two complementary attention mechanisms. Audio-Guided Visual Attention uses audio representations to identify informative visual cues, while Video-Guided Audio Attention uses visual representations to identify informative audio cues. Performing attention directly in hyperbolic space enables the model to capture nuanced relationships between speech characteristics and visual behavior that are indicative of hateful intent.
The attended audio-to-visual and visual-to-audio representations are combined through hyperbolic fusion, producing a unified multimodal representation that preserves the geometric structure learned in hyperbolic space. The fused representation is passed through a fully connected classification network, which predicts whether the input video contains hate speech or non-hate speech.
| Model | HateMM | MH-En | MH-Ch | ToxCMM | ||||
|---|---|---|---|---|---|---|---|---|
| Acc | F1 | Acc | F1 | Acc | F1 | Acc | F1 | |
| In-Distribution Evaluation | ||||||||
| A | .798 | .786 | .648 | .636 | .642 | .629 | .781 | .774 |
| V | .815 | .804 | .664 | .652 | .661 | .649 | .768 | .761 |
| CON | .836 | .826 | .708 | .695 | .701 | .689 | .807 | .797 |
| ECA | .861 | .848 | .748 | .735 | .755 | .742 | .848 | .835 |
| MA | .883 | .871 | .782 | .779 | .739 | .725 | .835 | .825 |
| VINAYAKA | .914 | .901 | .849 | .841 | .847 | .831 | .892 | .885 |
| Out-of-Distribution / Cross-Dataset Evaluation | ||||||||
| Training Dataset: HateMM (HMM) | ||||||||
| MM-HSD | .878 | .874 | .623 | .602 | .416 | .422 | .597 | .581 |
| MHC | .809 | .795 | .697 | .681 | .511 | .506 | .556 | .542 |
| TVLM | .833 | .821 | .769 | .761 | .497 | .480 | .632 | .641 |
| VINAYAKA | .914 | .901 | .808 | .789 | .563 | .551 | .724 | .712 |
| Training Dataset: MultiHateClip-En (MHCE) | ||||||||
| MM-HSD | .733 | .717 | .815 | .801 | .507 | .483 | .622 | .609 |
| MHC | .711 | .727 | .810 | .790 | .595 | .541 | .639 | .622 |
| TVLM | .755 | .729 | .833 | .812 | .520 | .503 | .641 | .622 |
| VINAYAKA | .797 | .786 | .849 | .841 | .626 | .619 | .715 | .706 |
| Training Dataset: MultiHateClip-Ch (MHCC) | ||||||||
| MM-HSD | .538 | .540 | .518 | .492 | .834 | .820 | .555 | .527 |
| MHC | .446 | .425 | .579 | .562 | .800 | .780 | .411 | .392 |
| TVLM | .528 | .515 | .299 | .330 | .609 | .576 | .527 | .508 |
| VINAYAKA | .566 | .557 | .619 | .616 | .847 | .831 | .621 | .612 |
| Training Dataset: ToxCMM | ||||||||
| MM-HSD | .602 | .581 | .672 | .657 | .410 | .390 | .828 | .819 |
| MHC | .596 | .577 | .649 | .633 | .555 | .534 | .789 | .771 |
| TVLM | .558 | .522 | .700 | .683 | .555 | .521 | .823 | .842 |
| VINAYAKA | .726 | .712 | .709 | .693 | .585 | .577 | .892 | .885 |
Table. Performance across In-domain and Cross-dataset Evaluation. Acc and F1 refer to Accuracy and Macro-Average F1 respectively. A = Audio-only, V = Video-only, CON = Concatenation, ECA = Euclidean Cross-Attention, MA = Möbius Addition, and VINAYAKA denotes the proposed approach.
@inproceedings{vinayaka2026,
title={Multilingual Audio-Visual Hate Speech Detection via Cross-Modal Fusion in Hyperbolic Space},
author={Bhavinkumar Vinodbhai Kuwar and Orchic Chetia Phukan and Rajesh Sharma},
booktitle={Interspeech},
year={2026}
}