The data set d6505c89-c43d-4c28-8c4f-7351a5fd5528 is based on Single-cell Atlas of the Transcriptome and Chromatin Accessibility in the Human Retina, Nature Genetics, 2026. There are 3177310 cells. However, in Data S1, I see 1775529 in Number of QC'd Cells column. Why does quality control discard 1.4 million cells? If there were low quality cells, wouldn't they have been removed before publication by the original authors? Retina is tough for quality control because it has some very transcriptionally quiet cell types.
Minimal filtering was applied to preserve biological heterogeneity.
It doesn't look like that, at least not for some data sets in the training set.
Outlier cells with expression levels exceeding 3 standard deviations above or below the dataset average were filtered.
That looks problematic for a couple of well-known abundant but transcriptionally quiet retinal cell types.
The data set d6505c89-c43d-4c28-8c4f-7351a5fd5528 is based on Single-cell Atlas of the Transcriptome and Chromatin Accessibility in the Human Retina, Nature Genetics, 2026. There are 3177310 cells. However, in Data S1, I see 1775529 in Number of QC'd Cells column. Why does quality control discard 1.4 million cells? If there were low quality cells, wouldn't they have been removed before publication by the original authors? Retina is tough for quality control because it has some very transcriptionally quiet cell types.
It doesn't look like that, at least not for some data sets in the training set.
That looks problematic for a couple of well-known abundant but transcriptionally quiet retinal cell types.