Embedding human judgment in crowd aggregation: a case study in biomedical citizen science
P. Michelucci*,
L. Onac,
L. Gusman,
M. Lane,
C. Nickerson,
S. Shaaban,
J. Lou,
S. Magaki,
L. Minaud,
M. Keiser and
B.N. Dugger*: corresponding author
Published on:
June 30, 2026
Abstract
Wisdom-of-crowds aggregation in citizen science relies on algorithmic processes to combine many non-expert contributions into expert-grade outputs [1]. When aggregation methods require large numbers of contributors, improvement efforts typically focus on refining the automation itself. We instead identified a specific capability gap within an aggregation algorithm and proposed a minimal human-in-the-loop (HITL) intervention targeted at that failure point. We applied this approach to Beta Catchers, a citizen-science project in which volunteers annotate Alzheimer-related lesions on whole-slide brain images to support studies of social and biological determinants of Alzheimer disease. A recent pilot study found that 30 volunteers per image were needed to achieve expert-level performance under the project's successive collapsing algorithm (SCA). We traced this requirement to a brittle aggregation step that relied solely on spatial proximity to determine whether nearby annotations referred to the same lesion. To address this limitation, we proposed a HITL task in which volunteers would determine whether candidate annotation pairs referred to the same lesion by viewing them in the context of the underlying image. To minimize the burden introduced by the HITL task, we further developed a confidence-gated triage mechanism in which empirically derived thresholds would automatically resolve high-confidence cases and route only lower-confidence cases for human judgment. Because errors in this low-confidence region appeared to account for much of the variance observed in the pilot study, we hypothesized that improving pairwise collapsing accuracy could reduce required contributors from 30 to approximately 10 per image. Applied to a 30-megatile pilot dataset, the triage mechanism reduced this burden to under two minutes of volunteer work per image, preserving the projected 3-fold throughput gain. This pattern of targeted HITL intervention may apply broadly to capability gaps in wisdom-of-crowds aggregation algorithms.
DOI: https://doi.org/10.22323/1.524.0020
How to cite
Metadata are provided both in
article format (very
similar to INSPIRE)
as this helps creating very compact bibliographies which
can be beneficial to authors and readers, and in
proceeding format which
is more detailed and complete.