Abstract

This thesis presents the results of investigation into methods of reducing bias in machine-learning-based content moderation. Using methods including regression models, decision trees, and transformer models, I plan to investigate how out-of-distribution data results in a decrease in model accuracy, which models are most resilient to these kinds of errors, and what other methods can be employed in combination with these models to otherwise improve out-of-distribution performance.

Content warning:

This project involves the analysis of data which contains instances of hate speech, discrimination, and other forms of offensive language directed toward minority groups and others. Some terminology is included in Chapter 4.1, Decision Tree Models as part of an analysis of model performance.

Advisor

Smith, Joseph

Department

Statistical and Data Sciences

Disciplines

Data Science

Keywords

content moderation, machine learning, OOD

Publication Date

2026

Degree Granted

Bachelor of Arts

Document Type

Senior Independent Study Thesis

Share

COinS
 

© Copyright 2026 Nathan J. Dieterich