VINAYAKA: Multilingual Audio-Visual Hate Speech Detection via Cross-Modal Fusion in Hyperbolic Space

Interspeech 2026

Bhavinkumar Vinodbhai Kuwar*1,2 Orchid Chetia Phukan*3 Rajesh Sharma1,4

1 Plaksha University, India

2 IIIT Delhi, India

3 NTHU, Taiwan

4 University of Tartu, Estonia

VINAYAKA Architecture

Figure 1. Overview of VINAYAKA. Audio representations extracted using WavLM and visual representations extracted using ImageBind are projected into hyperbolic space where cross-modal attention and hyperbolic fusion jointly model multilingual hateful intent.

Abstract

In this work, we investigate the robustness of audio–visual (AV) cues for Multilingual Audio–Visual Hate Speech Detection (M-AVHSD). We propose VINAYAKA, a novel framework for M-AVHSD that relies solely on AV cues for decision-making. Along with WavLM and ImageBind, it employs Cross-modal Fusion in Hyperbolic Space (CFHS), which models the effective alignment between paralinguistic and behavioral AV cues to better capture hateful intent in temporal audio signals and corresponding visual scenes. Empirical results demonstrate that the proposed framework achieves state-of-the-art performance in both in-distribution and out-of-distribution settings across cross-lingual and cross-dataset evaluations. Furthermore, the results show that the framework remains robust to ASR error propagation.

Key Contributions

Text-Free Detection

Detects hateful intent using only audio-visual cues without relying on transcriptions.

Hyperbolic Fusion

Introduces Cross-modal Fusion in Hyperbolic Space (CFHS) to model interactions between audio and visual behavioral signals.

Cross-Lingual Robustness

Generalizes across languages and datasets under challenging out-of-distribution conditions.

ASR Independent

Remains robust even when automatic speech recognition systems introduce errors.

Method

VINAYAKA (Multilingual Audio-Visual Hate Speech Detection via Cross-Modal Fusion in Hyperbolic Space) detects hateful intent directly from audio and visual signals without relying on textual transcriptions. The framework combines pretrained audio and visual representations with hyperbolic cross-modal reasoning to model interactions between paralinguistic speech cues and behavioral visual cues. By performing multimodal fusion in hyperbolic space, VINAYAKA captures hierarchical relationships between modalities and improves generalization across languages and datasets.

1

Feature Extraction

Raw audio waveforms are processed using WavLM, while video frames are encoded using the ImageBind Visual Encoder. These pretrained encoders generate rich semantic and behavioral representations for each modality. The resulting audio embedding E(a) and visual embedding E(v) capture complementary information relevant to multilingual hate speech detection.

2

1D-CNN and Temporal Aggregation

Audio and visual representations are independently passed through a lightweight 1D-CNN followed by max-pooling. This stage refines modality-specific patterns, suppresses noise, and aggregates temporal information. The pooled representations are flattened to produce compact feature vectors suitable for multimodal interaction.

3

Geometric Projection into Hyperbolic Space

The flattened Euclidean representations are projected into a shared hyperbolic space. Unlike Euclidean geometry, hyperbolic geometry naturally models hierarchical and complex relational structures. The projected embeddings form hyperbolic audio sequences and hyperbolic video sequences that provide a richer representation space for cross-modal reasoning.

4

Hyperbolic Cross-Attention

VINAYAKA performs bidirectional cross-modal reasoning through two complementary attention mechanisms. Audio-Guided Visual Attention uses audio representations to identify informative visual cues, while Video-Guided Audio Attention uses visual representations to identify informative audio cues. Performing attention directly in hyperbolic space enables the model to capture nuanced relationships between speech characteristics and visual behavior that are indicative of hateful intent.

5

Hyperbolic Fusion and Classification

The attended audio-to-visual and visual-to-audio representations are combined through hyperbolic fusion, producing a unified multimodal representation that preserves the geometric structure learned in hyperbolic space. The fused representation is passed through a fully connected classification network, which predicts whether the input video contains hate speech or non-hate speech.

Performance across In-domain and Cross-dataset Evaluation. Acc and F1 refer to Accuracy and Macro-Average F1 respectively.
Model HateMM MH-En MH-Ch ToxCMM
Acc F1 Acc F1 Acc F1 Acc F1
In-Distribution Evaluation
A .798.786 .648.636 .642.629 .781.774
V .815.804 .664.652 .661.649 .768.761
CON .836.826 .708.695 .701.689 .807.797
ECA .861.848 .748.735 .755.742 .848.835
MA .883.871 .782.779 .739.725 .835.825
VINAYAKA .914.901 .849.841 .847.831 .892.885
Out-of-Distribution / Cross-Dataset Evaluation
Training Dataset: HateMM (HMM)
MM-HSD .878.874 .623.602 .416.422 .597.581
MHC .809.795 .697.681 .511.506 .556.542
TVLM .833.821 .769.761 .497.480 .632.641
VINAYAKA .914.901 .808.789 .563.551 .724.712
Training Dataset: MultiHateClip-En (MHCE)
MM-HSD .733.717 .815.801 .507.483 .622.609
MHC .711.727 .810.790 .595.541 .639.622
TVLM .755.729 .833.812 .520.503 .641.622
VINAYAKA .797.786 .849.841 .626.619 .715.706
Training Dataset: MultiHateClip-Ch (MHCC)
MM-HSD .538.540 .518.492 .834.820 .555.527
MHC .446.425 .579.562 .800.780 .411.392
TVLM .528.515 .299.330 .609.576 .527.508
VINAYAKA .566.557 .619.616 .847.831 .621.612
Training Dataset: ToxCMM
MM-HSD .602.581 .672.657 .410.390 .828.819
MHC .596.577 .649.633 .555.534 .789.771
TVLM .558.522 .700.683 .555.521 .823.842
VINAYAKA .726.712 .709.693 .585.577 .892.885

Table. Performance across In-domain and Cross-dataset Evaluation. Acc and F1 refer to Accuracy and Macro-Average F1 respectively. A = Audio-only, V = Video-only, CON = Concatenation, ECA = Euclidean Cross-Attention, MA = Möbius Addition, and VINAYAKA denotes the proposed approach.

Citation

@inproceedings{vinayaka2026,
    title={Multilingual Audio-Visual Hate Speech Detection via Cross-Modal Fusion in Hyperbolic Space},
    author={Bhavinkumar Vinodbhai Kuwar and Orchic Chetia Phukan and Rajesh Sharma},
    booktitle={Interspeech},
    year={2026}
}
        

Contact

For questions regarding the paper, please contact

bhavinkumar24212@iiitd.ac.in