ef08d55de458d633b75e54be97bf00c6a9377eba
max
  Mon Sep 21 05:19:50 2026 -0700
hg38 episignatures: link PubMed as "Lastname Year" instead of the bare PMID

The Studies table on methaDory.html and the locus table on epigenCentral.html
linked out to PubMed with the PMID digits as the clickable text, which reads
oddly next to a plain-text study/gene identifier. Both now show the first
author's last name and the publication year instead, from a new checked-in
lookup, scripts/episignatures/pubAuthorYear.tsv, built from PubMed esummary
(and Crossref for the one DOI-only reference). makeHtmlTables.py and
epigenCentralToBed.py errAbort if an id used by the data has no entry, so a
future refresh can't silently regress to showing raw PMIDs again.

The lookup's year is PubMed's own citable pubdate, which for a handful of
entries differs by one from the StudyID naming used elsewhere on the page
(assigned by the source labs, sometimes from an epub-ahead-of-print date);
that's expected, not a mismatch to fix.

The two hand-written References sections at the bottom of each page keep the
usual "PMID: <number>" convention used across every other track's HTML.

diff --git src/hg/makeDb/doc/hg38/episignatures.txt src/hg/makeDb/doc/hg38/episignatures.txt
index 03df792690b..826e2a9c354 100644
--- src/hg/makeDb/doc/hg38/episignatures.txt
+++ src/hg/makeDb/doc/hg38/episignatures.txt
@@ -85,30 +85,36 @@
 # filterValues.loci, filterValues.studies) are generated by the build into
 # methaDoryFilters.ra. When the source data is refreshed, paste those three
 # pairs of lines back into human/hg38/episignatures.ra rather than editing them
 # by hand.
 
 # The two tables on the description page are generated the same way:
 python3 ~/kent/src/hg/makeDb/scripts/episignatures/makeHtmlTables.py \
     studies studySummary.tsv > studyTable.html
 python3 ~/kent/src/hg/makeDb/scripts/episignatures/makeHtmlTables.py \
     loci locusSummary.tsv > locusTable.html
 # and pasted into human/hg38/methaDory.html between the
 # "<!-- BEGIN generated ... -->" and "<!-- END generated -->" markers under the
 # "Studies" and "Genes and loci" headings. The counts quoted in the Description
 # and Methods paragraphs of that page have to be updated by hand at the same
 # time.
+#
+# Both tables link out to PubMed with a "Lastname Year" label rather than the bare
+# PMID, looked up from scripts/episignatures/pubAuthorYear.tsv (also used by
+# epigenCentralToBed.py for the EpigenCentral locus table). On a refresh, add an
+# entry for every new PMID before regenerating the tables; the scripts errAbort on
+# an id with no entry rather than falling back to the raw number.
 
 # Colour: five classes, from the strongest effect at the site. Warm = hyper,
 # cool = hypo, darker = |delta-beta| >= 0.10, purple = conflicting. Three things
 # decided this, all measured on the data rather than assumed:
 #
 #   - A linear ramp on delta-beta would be useless. The distribution is heavily
 #     right-skewed: median 0.110, p75 0.159, p90 0.232, p99 0.416, max 0.835,
 #     and 43% of sites sit in 0.05-0.10. Nearly everything would land in the
 #     bottom fifth of the ramp.
 #
 #   - The low end is a reporting artefact, not biology. Per study, the minimum
 #     |delta-beta| reveals the cut-off each paper applied before publishing its
 #     site list: 4 studies cut at 0.20 (Velasco2021 among them, whose median is
 #     therefore 0.254 against Levy2022's 0.076), 25 at ~0.10, 33 at ~0.05, and
 #     12 applied none. 98.1% of rows come from a study that cut at >= 0.05.