整個研究的「語言」。每個術語給 一句白話 → 直覺類比 → (折疊)嚴謹定義 + 常見誤解;統計方法的詞另附 worked example(怎麼算)。其他頁的術語首次出現都連回這裡。
這頁是全站名詞的單一真相來源(SoT)。任何頁看到不懂的術語,點它連回這裡的對應卡片。讀到夠就停——L0 標題行 + 直覺類比通常就夠;要嚴謹定義/公式再展開 ▸。
▸ 給 PI:本研究最關鍵的語言是 HP-axis(project-specific somatic-conditioned)vs ALLELE-axis(易混入 germline baseline)、「observed association ≠ 能判別」、以及 excess-over-null(不能比 raw rate)——這三個定義決定了所有結論能不能成立。
🧮 = 附 worked example · 🧩 = 高認知負荷 · 完整 92 詞 SoT:InterSubMod/docs/explain/_planning/02_名詞表_glossary.md
「這條 read 來自哪一條染色體版本」這件事的詞彙。
常見誤解:把 HP1/HP2 當「第幾號染色體」;它是「同一對同源染色體的父本 vs 母本版本」。關聯:HP-tag · HP1-1 · Phasing
HP:i:)— BAM 裡每條 read 的單倍型標籤longphase-S(paired)下 HP=1/2 為 germline;somatic haplotagging 6-state 另有 HP=11(HP1-1)/21(HP2-1)/HP3/HP0。常見誤解:把 HP=11 當第三條染色體或 copy number 標記。
常見誤解:(1) 誤為第三單倍型;(2) 誤為 copy loss / CN tag;(3) 誤為一定能乾淨偵測(依賴 longphase-S,LOH 區反而最難 phase)。圖見 ISM 頁。
常見誤解:以為到處可 phase;LOH 區 germline het SNP 消失 → 失去錨點(這正是 methyl-assisted phasing 想救的白地)。
常見誤解:以為「救援 unphase」到處可行;多數 unphase read 落在 LOH/imprinting 區(germline SNP 稀疏)→ 無本地錨點建甲基參考 → 機制恰在最需要處失效(chicken-egg)。
常見誤解:SEQC2 LOH annotation 或 HP imbalance 本身不能區分 deletion/CN-loss、copy-neutral LOH、保留哪個 parental copy 或其他機制;CN 尚未整合。也不可把此 phasing signature 當「甲基化雙峰」(HPFineNGroups 曾被誤解,實為 phasing×allele occupancy count)。⚠ 帶 by-construction circularity。
移除 self-phasing 後大量 TO-mode TP LOH 消失。決定性負控需 --germline-hp-only flag 關 somatic tagging 重觀察(R-SELFREF);未跑前 phasing 脊柱證據封頂 Grade B+ 非 Grade A。
常見誤解:因名字像「HP 細分 N 個甲基 group」而誤為甲基雙峰;C++ 原始碼(LabelTest.cpp hp_to_fine_labels)確認為 occupancy count。→ feature name 必查源碼的經典案例。
somatic / germline、LOH、copy number — 用來界定 ASM 是否與 normal-anchored cis-association candidate 相容,以及仍有哪些 confound;這類 association 不證明因果或 tumor-acquired。
paired tumor-normal calling(ClairS)通常改善 somatic/germline 分離,但仍不完美,必須以 truth benchmark 驗證;ClairS-TO(tumor-only)分離限制更大。常見誤解:以為 caller 報出都是 somatic。關聯:TP/FP · germline-het 對照
區分重要,因為 ASM 的 copy confound 要看實際 CN,cnLOH 與 CN-loss 對 dosage 假象的影響不同。
文獻:apparent ASM 有 82–92% 可被 copy 數解釋(Martin-Trujillo 2017)→ 本專案以 normal-anchored cis heuristic 標記與部分 observed copy pattern 不相容的候選;它不是完整 CN 校正,也不證明 copy confound 已移除。
● 鐵則:ASM「存在」≠「能判別 TP/FP」。本研究 strong-ASM 反而在 FP 富集(OR=0.194)。
HCC1395 主 benchmark 真值 = SEQC2 v1.2.1 的 39,447 somatic SNV。ISM 多輪研究問「甲基特徵能否分 TP/FP」→ 已測 formulation 的答案為 negative;若有 materially new、經 pre-decision audit 的假說,仍可重開驗證。
ISM 讀的訊號本身。
ISM 對每位點 ±1000bp 視窗內的 CpG 建 read×CpG 矩陣。
⚠ ISM 目前 5mC-only(max-collapse 把 5mC+5hmC 議題:盤點建議分軌處理,dup-bug 待修)。
⚠ Δβ(甲基率差,有方向)≠ Δ(距離差,無方向)。見 ISM 頁圖。
ISM 的 MethylationParser 解這兩個 tag,CIGAR→ref 定位 CpG。⚠ MM/ML 解析目前零單元測試(盤點稽核 flag)。
這是 ISM 站「軸 C(between-read 距離)」對比業界「軸 A(per-position 率差)」的根本差異。
方法學的統計詞,每個附「怎麼從數字算出來」的 worked example(輸入為示意數字,算式取自源碼)。
DistanceMatrix.cpp:23)。
--nan-distance-strategy SKIP 標記為無效對並排除。DistanceConfig 的 internal initializer 5 與 legacy/experimental MAX_DIST=1.0 會被 current Config 的 3/SKIP 覆寫,不是無參數生產行為。有效 CLI 在未指定 --distance-metric 時使用 NHD;Config 的 BERNOULLI initializer 會被參數解析器清空,只有明示請求 BERNOULLI 才會走 ML 機率加權路徑。
StructureTest.cpp:150)。
src/core/RegionProcessor.cpp 的 structure_config.n_permutations 與 include/core/LabelTest.hpp 均為 999)重算 F,看觀測 F=20 多極端 → p(現行實作分母為 999+1,解析度下限 0.001)。(SS 為示意;實際由距離矩陣算)
常見誤解:把它當「per-CpG 率差」或「unsupervised cluster 存在」檢定。它只測指定 label 與距離 variation/centroid separation 在置換 null 下的 association,且 dispersion 差異也能影響結果,必須與 PERMDISP 合讀;不能證明 cluster truth、cellular group 或因果。
MathUtils.hpp:85,min(r,c)−1=1)。
MathUtils.cpp::cramers_v)才採信;raw V 與旗標分開保存,RegionProcessor 的 summary 才投影為 reliable ? raw_v : 0。
最容易誤讀的軸別定義,與上下游工具。
● 鐵則 2:宣稱 somatic ASM 必附 germline-het negative control。
啟發式判準 |d_cis| > 1.8×|d_drift| + 0.02。例 BRCA2 d_within 小、d_copy 大 → 與 observed copy/dosage pattern 相容;chr17/TBC1D16 d_within=0.142, p=0.001 → cis-compatible/copy-screen-negative candidate。此 project-specific heuristic 僅 screen 部分 observed confound,不證明 copy、germline 或 drift 已移除,也不證明 causal cis 或 tumor-acquired。● 單一樣本、僅 HCC1395 有 matched normal