Features or Spurious Artifacts? Data-centric Baselines for Fair and Robust Hate Speech Detection

Alan Ramponi; Sara Tonelli

doi:10.18653/v1/2022.naacl-main.221

ScienceGate Book Chapters

JOURNAL ARTICLE

Features or Spurious Artifacts? Data-centric Baselines for Fair and Robust Hate Speech Detection

Alan Ramponi Sara Tonelli

Year: 2022 Journal: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies

DOI: 10.18653/v1/2022.naacl-main.221

Get Full-Text PDF Get Analytical Report

Abstract

Avoiding to rely on dataset artifacts to predict hate speech is at the cornerstone of robust and fair hate speech detection. In this paper we critically analyze lexical biases in hate speech detection via a cross-platform study, disentangling various types of spurious and authentic artifacts and analyzing their impact on out-of-distribution fairness and robustness. We experiment with existing approaches and propose simple yet surprisingly effective data-centric baselines. Our results on English data across four platforms show that distinct spurious artifacts require different treatments to ultimately attain both robustness and fairness in hate speech detection. To encourage research in this direction, we release all baseline models and the code to compute artifacts, pointing it out as a complementary and necessary addition to the data statements practice.

Keywords:

Spurious relationship Robustness (evolution) Computer science Voice activity detection Speech recognition Artificial intelligence Machine learning Natural language processing Speech processing

Metrics

Cited By

2.00

FWCI (Field Weighted Citation Impact)

Refs

0.87

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Hate Speech and Cyberbullying Detection

Physical Sciences → Computer Science → Artificial Intelligence

Adversarial Robustness in Machine Learning

Physical Sciences → Computer Science → Artificial Intelligence

Features or Spurious Artifacts? Data-centric Baselines for Fair and Robust Hate Speech Detection

Abstract

Metrics

Citation History

Topics

Related Documents

Robust Hate Speech Detection via Mitigating Spurious Correlations

Improving Counterfactual Generation for Fair Hate Speech Detection

Robust Implementation of a Hate Speech Detection System

Towards more robust hate speech detection: using social context and user data

Hate Speech Detection Using Textual and User Features