Improving Label Quality by Jointly Modeling Items and Annotators

Jan 1, 2021·

Tharindu Cyril Weerasooriya

Alexander G. Ororbia II

Christopher M. Homan

· 0 min read

Abstract

We propose a fully Bayesian framework for learning ground truth labels from noisy annotators. Our framework ensures scalability by factoring a generative, Bayesian soft clustering model over label distributions into the classic David and Skene joint annotator-data model. Earlier research along these lines has neither fully incorporated label distributions nor explored clustering by annotators only or data only. Our framework incorporates all of these properties as: (1) a graphical model designed to provide better ground truth estimates of annotator responses as input to any black box supervised learning algorithm, and (2) a standalone neural model whose internal structure captures many of the properties of the graphical model. We conduct supervised learning experiments using both models and compare them to the performance of one baseline and a state-of-the-art model.

Type

Publication

CoRR

Last updated on Jan 1, 2021

Label Distribution Learning Bayesian Methods Neural Networks Annotator Quality Machine Learning

← Improving Label Quality by Joint Probabilistic Modeling of Items and Annotators Jan 1, 2022

Neighborhood-based pooling for population-level label distribution learning Jan 1, 2020 →