Exploring Intensities of Hate Speech on Social Media: A Case Study on Explaining Multilingual Models with XAI (JKU Visual Data Science Lab)

Jul 14, 2023 #Algorithms, #Policies

In order to present a more accurate spectrum of severity and surmount the constraints of seeing hate speech as a binary task (as typical in sentiment analysis), we classify hate speech into four intensities: no hate, intimidation, offense or discrimination, and promotion of violence. For this, we first involve 31 users in annotating a dataset in English and German. To promote interpretability and transparency, we integrate our ML system in a dashboard provided with explainable AI (XAI). By performing a case study with 40 non-experts moderators, we evaluated the efficacy of the proposed XAI dashboard in supporting content moderation. Our results suggest that assessing hate intensities is important for content moderators, as these can be related to specific penalties. Similarly, XAI seems to be a promising method to improve ML trustworthiness, by this, facilitating moderators’ well-informed decision-making.

https://jku-vds-lab.at/publications/2023_ditox_hate_speech_xai/

Exploring Intensities of Hate Speech on Social Media: A Case Study on Explaining Multilingual Models with XAI (JKU Visual Data Science Lab)

Like this:

Leave a Reply Cancel reply

LATEST NEWS

New on preventhate.org, 12 July 2026 (Policyinstitute.net)

UNESCO launches issue brief on Media and Information Literacy to counter hate speech in the digital age (UNESCO)

Five lessons from the No Hate Speech Week: what we heard, what we learned, what comes next (Council of Europe)

Hate speech levels across Europe alarming, stronger action needed (Council of Europe)

Soft Security Resources: Press Articles, Documents, and Recordings on Countering Extremism, Hate Speech, and False Information – December 2025 (II/II)

TAGS

preventhate.org | Policyinstitute.net

Exploring Intensities of Hate Speech on Social Media: A Case Study on Explaining Multilingual Models with XAI (JKU Visual Data Science Lab)

Share this:

Like this:

Leave a Reply Cancel reply

New on preventhate.org, 12 July 2026 (Policyinstitute.net)

UNESCO launches issue brief on Media and Information Literacy to counter hate speech in the digital age (UNESCO)

Five lessons from the No Hate Speech Week: what we heard, what we learned, what comes next (Council of Europe)

Hate speech levels across Europe alarming, stronger action needed (Council of Europe)

Soft Security Resources: Press Articles, Documents, and Recordings on Countering Extremism, Hate Speech, and False Information – December 2025 (II/II)

TAGS