54f465e316441236c128491e53c0e5464c0e672c lrnassar Tue Sep 29 11:40:49 2026 -0700 Fiber-seq: bring the fiberSeqTrackDb.py docstrings back in line with the three-value sampleClass, and stop the Compendium intro from keeping a running tally of where the lymphoblastoid lines came from. readSamples still said the five GM lines were common cell lines and writeMetadata still said there were two classes, both a few lines from the SAMPLE_CLASS_COLORS entry that added the third. writeMetadata also still had "Common Cell Line" in the old Title Case, which yesterday's rename missed because the string is wrapped across two source lines and a line oriented sed cannot see it; worth remembering for the next rename. The intro sentence had grown a breakdown that did not add up, 20 HPRC plus 5 rare disease against 27 lymphoblastoid lines, leaving GM12878 and HG002 unaccounted; it now points at the Sample class filter instead of counting, since the lab has said more rare disease samples are coming and the tally would go stale again. While in there, the claim that accession order keeps each class together is softened to what the data actually does: the common cell lines fall in two runs either side of the HPRC block. Generated output is unchanged; the .ra, the metadata TSV and the colors JSON all regenerate byte identical. Caught by Claude review of 10f0f6d516. refs #36210 diff --git src/hg/makeDb/scripts/fiberSeq/fiberSeqTrackDb.py src/hg/makeDb/scripts/fiberSeq/fiberSeqTrackDb.py index 8390e926b88..de26f2f7591 100755 --- src/hg/makeDb/scripts/fiberSeq/fiberSeqTrackDb.py +++ src/hg/makeDb/scripts/fiberSeq/fiberSeqTrackDb.py @@ -70,72 +70,75 @@ "Human Pangenome Reference Consortium. Rare disease samples are cases " "consented to broad genomic data sharing") # Sample class swatches, shown next to that facet's checkboxes. # Okabe-Ito colors for the swatches. SAMPLE_CLASS_COLORS = { "HPRC": "#0072B2", "Common cell line": "#D55E00", "Rare disease sample": "#009E73", } def readSamples(path): """Read fiberSeqSamples.tsv into a list of dicts, in file order. - sampleClass comes from the lab's own sample sheet, not from the cell type: - five of the lymphoblastoid lines are common cell lines rather than HPRC - samples, so there is nothing in the cell type that tells the two apart.""" + sampleClass comes from the lab's own sample sheet, not from the cell type. + The 27 lymphoblastoid lines split three ways, 20 HPRC and 5 rare disease + samples and 2 common cell lines, and nothing in the cell type tells them + apart. Expect the rare disease group to grow; the lab is sending more.""" samples = [] with open(path) as f: for line in f: if line.startswith("#") or not line.strip(): continue fields = line.rstrip("\n").split("\t") if len(fields) < 5: sys.exit("bad sample line, want 5 fields: %s" % line.rstrip()) acc, sample, cellType, _hash, sampleClass = fields[:5] samples.append({ "accession": acc, "sample": sample, "cellType": cellType, "sampleClass": sampleClass, }) if not samples: sys.exit("no samples read from %s" % path) return samples def writeMetadata(path, samples): """The sample table shown on the track UI page. A plain column name gets facet checkboxes, a leading underscore means searchable and sortable but not faceted. Accession is the primaryKey but sits last, since it is the least interesting thing about a sample. Nothing requires the primaryKey to come first: facetedComposite.js only checks that the column exists, and every use of it - is by name. It is still the default sort, because accession order keeps the - common cell lines together and then the HPRC samples together, which sample - name in alphabetical order would scatter. - - Sample class is the one faceted column. Its two values, HPRC and Common - Cell Line, each cover many samples, which is what a facet needs: - facetedComposite.js only offers a value that occurs more than once, since a - checkbox matching a single row is just a slow search box. The other three - are underscored for that reason. Sample and Accession are unique per row by - definition, and 12 of the 14 cell types are a single sample, so as a facet - cell type drew two checkboxes and left 12 samples unreachable. + is by name. It is still the default sort, because accession order mostly + groups samples of a kind together, which sample name in alphabetical order + would scatter. Only mostly: the accessions were assigned as the lab + produced the data, so the common cell lines fall in two runs either side of + the HPRC block. + + Sample class is the one faceted column. Its three values, HPRC, Common cell + line and Rare disease sample, each cover many samples, which is what a facet + needs: facetedComposite.js only offers a value that occurs more than once, + since a checkbox matching a single row is just a slow search box. The other + three are underscored for that reason. Sample and Accession are unique per + row by definition, and 12 of the 14 cell types are a single sample, so as a + facet cell type drew two checkboxes and left 12 samples unreachable. Names are underscore separated rather than camelCase: the header is rendered by toTitleStyle() in facetedComposite.js, which turns an underscore into a space but does not split camelCase, so "sampleClass" would have read "sampleClass" in the table. A literal space cannot be used instead, because the saved sort order is a space separated list of column names and the submit code drops any name containing whitespace.""" with open(path, "w") as f: # A header cell may carry a longer description after a "|", which the # track UI shows behind an info icon on the column heading. f.write("_Sample\t_Cell_type\tSample_class|%s\tAccession\n" % SAMPLE_CLASS_DESCRIPTION) for s in samples: f.write("%s\t%s\t%s\t%s\n" % (s["sample"], s["cellType"], s["sampleClass"], s["accession"]))