Skip to content

Quality Control Removes Over a Million Cells From One Dataset #63

Description

@DarioS

The data set d6505c89-c43d-4c28-8c4f-7351a5fd5528 is based on Single-cell Atlas of the Transcriptome and Chromatin Accessibility in the Human Retina, Nature Genetics, 2026. There are 3177310 cells. However, in Data S1, I see 1775529 in Number of QC'd Cells column. Why does quality control discard 1.4 million cells? If there were low quality cells, wouldn't they have been removed before publication by the original authors? Retina is tough for quality control because it has some very transcriptionally quiet cell types.

Minimal filtering was applied to preserve biological heterogeneity.

It doesn't look like that, at least not for some data sets in the training set.

Outlier cells with expression levels exceeding 3 standard deviations above or below the dataset average were filtered.

That looks problematic for a couple of well-known abundant but transcriptionally quiet retinal cell types.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions