解釋中心
解釋中心 · 方法本體 · ism-core

ISM — read-level 甲基化結構分析引擎

這頁回答:ISM 是什麼、坐在整條分析流程的哪裡、吃什麼吐什麼、跟業界做法差在哪、最容易誤讀的幾個點。讀到夠就停——L0 重點平鋪、細節折在 ▸ 內、重要概念都配圖。

L0 一句話結論

ISM(InterSubMod)= 對每個 somatic 突變位點,把周圍 reads 的甲基化模式做成 read×read 距離矩陣 → 階層聚類 → 指定 label 的 PERMANOVA,量「read-distance variation/centroid separation 是否與指定單倍型或 somatic label 關聯」。它不是 variant filter;PERMANOVA 也不驗證 unsupervised cluster、cellular group 或因果。

▸ 給 PI 一句話:ISM 描述 between-molecule(分子間)距離與 label association;project-specific normal-anchored cis heuristic 只標記 cis-compatible/copy-screen-negative candidate,不能宣稱 copy/germline/drift confound 已移除。

199 欄
frozen release baseline ddd8909a 的 significance_summary.csv source header● P
0.34%
全基因組 strong-ASM 稀有率(hypo 44%/hyper 56% 無偏好)★3
−0.122
具名 BRCA2、既定 HP-axis 的 observed scope-bound Δβ;非 causal effect ● P
TESTED NEG.
已測甲基→TP/FP formulation:OR=0.194, p=1.8e-28;LOSO 100% 循環 ● L4

怎麼讀:名詞建直覺 → 作用位置看流程圖 → 為何/發現/結論看邏輯 → 難點拆解是最會卡的 4 個點(都配圖)→ 數據/溯源給嚴謹引用。

● 全站 4 條鐵則(讀任何一頁都要記得)
  1. 「具名 scope 有 allele-associated methylation」≠「能判別 TP/FP」——兩個 claim 永遠分開講。
  2. HP-axis(HP1 vs HP1-1)是 project-specific somatic-conditioned comparison,可減少部分 context confound;ALLELE-axis(ALT vs REF)易混入 germline het baseline。兩者都不是 confound-free/causal 軸,宣稱 somatic ASM 必附 germline-het 對照。
  3. 絕對 AUC 在甲基-phasing 場景系統性膨脹(germline het null median 0.974);引相對 null 或 held-out rescue rate,不引絕對 AUC。
  4. 甲基→TP/FP filter 的已測 formulation 為 concluded negative;只有 materially new hypothesis + pre-decision audit 才可重開。同 haplotype 內 T3 亦只對已測設計為 negative。

§1 名詞地基 N-DEF

這頁會用到的關鍵術語,每個給「一句白話 → 直覺類比 → (折疊)嚴謹定義 + 圖 + 來源」。canonical home 之後移到「背景名詞地基頁」,本頁先就地定義。

Δ(距離差)vs Δβ(甲基率差) — 兩個完全不同的量 ● 最常混淆
直覺:Δ 描述指定兩組 reads 的甲基模式距離差(無方向、單位是距離);Δβ 描述某 CpG 上觀測甲基化比例差(有方向、單位 0–1)。兩者都受既定分組與 scope 限制,不證明 unsupervised cluster truth 或 causal effect。
圖:同一位點的兩個不同量
Δ 距離空間與 Δβ 甲基率空間 對照 read-level 距離差與聚合甲基率差,說明兩者的定義與不可互換性。 距離空間 → Δ(無方向) read×read 距離 群內近 群間遠 Δ = d_between − d_within 例 HP-axis Δ=0.15 · 描述指定組距離差,不證明群體真值 甲基率空間 → Δβ(有方向) β HP1(高) HP1-1(低) Δβ = β(HP1-1) − β(HP1) = −0.122 有正負 · BRCA2 既定 scope 的 observed 差值,非 causal effect
左 = 距離空間(Δ):描述指定 label 組間與組內 read-distance 差,無方向。右 = 甲基率空間(Δβ):描述既定 scope 的 observed per-CpG 比例差,有正負方向。兩者都不是 unsupervised cluster truth、cellular group 或因果證據。

嚴謹定義:Δ = read-read 距離空間 d_between − d_within(單位 NHD);Δβ = 每 CpG β_A − β_B(有正負)。來源:盤點線1 high_cog_load。

HP1-1 — somatic 子單倍型,不是第三條染色體 ● 易誤讀
直覺:你有兩條染色體(父=HP1、母=HP2)。腫瘤在父那條上長出一個體細胞突變 → 帶這突變的 reads 被 longphase-S 標成 HP1-1。它仍是 HP1 那條,只是「帶了後天突變的那一份」。
圖:HP1 如何分出帶突變的 HP1-1
HP1、HP2 與 somatic-supporting HP1-1 read family 示意 germline haplotype 與帶 somatic ALT 的 read family;HP1-1 不是第三個 haplotype,也不等於已確認細胞 clone。 正常:兩條 germline 染色體 HP1(父) HP2(母) HP1 長 somatic ALT 腫瘤:HP1 分出帶突變的 HP1-1 HP1(無突變 reads) HP1-1(同染色體 + somatic ALT★) ★ HP2(母) HP-axisHP1 vs HP1-1部分 context 固定 ▸ HP-axis 比同 haplotype label 的「無突變 vs 有突變」reads;可減少部分 context confound,但非 confound-free 或 causal 軸。 ▸ 對比:ALLELE-axis(ALT vs REF)會把 germline 基線等位甲基差混進來 → 被 confound(鐵則 2)。 來源:HaplotagStrategy.cpp:505-516(盤點線 10)
HP1-1 = HP1 germline 鏈上帶 somatic ALT 的 reads(非第三 haplotype)。HP-axis 是本專案的 somatic-conditioned comparison,可減少部分 haplotype context confound;不證明 copy、germline 或 drift 已完全控制。
PERMANOVA — 置換檢定 label-associated centroid separation
直覺:把 reads 的 HP 標籤隨機洗牌很多次,檢查觀測標籤下的 centroid separation 是否與指定 null 不相容;仍需同看 PERMDISP 與設計限制,不能直接叫分群真值。

嚴謹定義:以距離矩陣算指定 label 的 pseudo-F =(between-label variation/within-label variation),再用標籤置換建 null 算 p。生產 pipeline 實跑 999 次置換(src/core/RegionProcessor.cpp 的 structure_config.n_permutations 與 include/core/LabelTest.hpp 均為 999)→ 現行實作分母為 999+1,p 解析度下限 = 0.001。結果須與 PERMDISP 合讀;顯著可與 centroid 或 dispersion 差異相容,不證明 unsupervised cluster、cellular group 或因果。● 源碼實測

算式 worked example · pseudo-F pseudo-F = (SS_between / (k−1)) / (SS_within / (n−k))(StructureTest.cpp:150)。
k=2 群, n=6:假設 SS_between=0.50, SS_within=0.10
→ F = (0.50 / 1) / (0.10 / 4) = 0.50 / 0.025 = 20
F 大 = 指定 HP label 下 between-label variation 相對較大。再把 HP 標籤洗牌 999 次重算 F,看觀測 F=20 在 null 分布多極端 → 得 p。(輸入為示意;實際 SS 由距離矩陣算;仍需 PERMDISP)
normal-anchored cis heuristic — 標記 cis-compatible/copy-screen-negative 候選
直覺:某 haplotype copy 變多,平均甲基可能因份數改變而偏移。matched-normal heuristic 只標記 cis-compatible/copy-screen-negative candidate,未被驗證為 causal test。

判準 |d_cis| > 1.8×|d_drift| + 0.02(scripts/34:149);copy-pattern screen |d_within| < 0.5×|d_HP|(scripts/37:64)。文獻:apparent ASM 有 82–92% 可被 copy 數解釋(Martin-Trujillo 2017)。此 project-specific heuristic 僅標記部分 observed confound,不證明 copy/germline/drift 已移除,也不證明 causal cis 或 tumor-acquired。● 僅 HCC1395 完整 (圖見 §5 難點 3)

Cramér's V reliability gate 與「487 association candidates」
直覺:cluster × label 表格太稀疏時,Cramér's V 關聯強度不可信。ISM 用 Cochran 規則(至少 80% 格的期望值 ≥5)另設 cramers_v_reliable 旗標;MathUtils 仍回傳 raw V。只有彙總欄位投影在不可靠時寫 0,這不判定 cluster 或 label 真值。

summary CramersV = reliable ? raw_v : 0(MathUtils.cpp::cramers_v、RegionProcessor.cpp 的 summary projection)。per-region significance.json 與 legacy GlobalTest::passed_gate 仍使用 raw V,不能把這三個口徑混為同一 gate。762 個 MISSED 中 487 個只描述歷史分析中的「summary V=0、指定 label 的 PERMANOVA 顯著」候選:須與 PERMDISP 合讀,不能升格為 unsupervised cluster、cellular group 或因果。🔵 S ● 不可寫成 filter (圖見 §5 難點 4)

算式 worked example · Cramér's V 2×2 時 V = √(χ² / n)(MathUtils.hpp:85,min(r,c)−1=1)。
表 [[30, 8], [6, 25]], n=69 → χ² ≈ 24.3 → V = √(24.3 / 69) ≈ 0.59
V∈[0,1],越大關聯越強。可靠性:需 ≥80% 格的期望值 ≥5(Cochran;GlobalTest.cpp:111-119)才採信,否則 gate 成 0。(示意表)

§2 作用位置:ISM 坐在流程的哪裡 N-ROLE

ISM 是一支核心 C++ binary,夾在上游 ONT 工具與下游 Python 後處理之間。

上游 ONT 工具(Dorado MM/ML · longphase-S HP tag) → 【 ISM C++ binary 】 → 下游 Python(observed Δβ · cis heuristic · copy-pattern partition)

輸入 / 輸出(吃什麼、吐什麼)

方向項目說明
輸入Tumor BAM含甲基 tag MM/ML + 單倍型 HP tag(1/2/11/21/33)
Normal BAM(可選)project-specific normal-anchored cis heuristic 需要
Reference FASTAhg38,CIGAR→ref 定位 CpG 用
Somatic SNV VCFClairS-TO;決定要分析哪些位點
輸出significance_summary.csv(frozen baseline source header:199 欄)全位點滙總主表;歷史 run 欄數不同,必須用欄名並記 producer commit ●
per-region 資料夾clustering/ · distance/ · methylation/ · reads/ · significance.json
下游量(observed Δβ / cis heuristic / copy-pattern partition)在 Python 層,不在 binary 內;不等於 validated causal test

五階段內部流程

ISM 五階段分析管線 從讀取與甲基矩陣,經 read-read 距離、分群與統計檢定,到輸出結果的流程。 ISM 核心五階段(模組名/常數取自 src/core/*,可 grep file:line) 每個 somatic SNV 由 OpenMP 獨立並行 · 不含上游 ONT 工具與下游 Python · 本圖 schematic 無分析數值 ① 輸入 + Read 擷取每 SNV 獨立並行 Somatic VCFSNV 位點 取視窗 ±1000bp 抓 read · BamReaderMAPQ≥20·len≥1000 ReadParser 過濾 HP tag1/2/1-1/2-1/3 ② 甲基矩陣建構MM/ML → 矩陣 · 僅 5mC 解 MM/ML5mC-only CIGAR→ref 定位 Read×CpG 矩陣MatrixBuilder 二值化 0.8/0.2 ③ Read-to-read 距離矩陣read×read · 無 imputation effective CLI 共同 CpG ≥3 距離 NHD/BERNOULLI N×N 矩陣 無效對→SKIP ④ UPGMA 階層聚類silhouette 選 k UPGMA linkage 選 k · silhouette 群標籤 TreeCutter ⑤ 顯著性檢定套件 → 輸出四路檢定 → current summary.csv 199 欄 GlobalTestFisher + CramérV gate StructureTestPERMANOVA 999-perm LabelTestHP/allele 置換 PerCpgAsm逐CpG Fisher+FDR 判定類別4 值 + Significant 輸出current summary 199 欄 ▣ label-flip:HP1-1 ≡ HP1 germline 鏈上帶 somatic ALT 的子單倍型(非第三 haplotype) ⚠ 預設 germline_hp_only=off → HP1-1 併入 HP-family 1;somatic 子單倍型軸需 hp_fine 4 群;cis heuristic/copy-pattern partition 在下游 Python;單樣本 HCC1395/★3
圖:ISM 五階段。① 輸入+取 read → ② 甲基矩陣 → ③ read×read 距離 → ④ UPGMA 聚類 → ⑤ 四路顯著性檢定。
※ 改編自 docs/paper_focus/04_figures/fig1_ism_5stage.svg;frozen release baseline ddd8909a 的 source header 為 199 欄,歷史 run 不得套用此固定欄位位置。
記住這個ISM 在流程裡的位置 = 把「上游給的甲基+單倍型 reads」轉成「每位點的 read-distance 與 label-association 統計」,丟給下游 Python 評估是否與 project-specific normal-anchored cis-compatible candidate 相容;這不證明 confound 已移除、因果或 tumor-acquired。它不碰 variant 對錯。

§3 為何做 / 發現什麼 / 結論

研究邏輯三幕。

為何做(動機 + gap)

ONT 長讀能在單一 read 上同時觀察甲基修飾(MM/ML)與單倍型(HP tag)。ISM 探索癌症 somatic sub-haplotype 維度,並加入 project-specific matched-normal copy-pattern screen;它不宣稱完整移除 copy confound。與現有工具的 novelty/比較優勢仍為 UNVERIFIED,待具日期的 systematic prior-art search 與 head-to-head benchmark。

發現什麼(三點,由重到輕)

發現證據 / 數字意涵
●甲基當 TP/FP filter:已測 formulation concluded negative多個 convergent tests:LOSO circularity · AUC=0.505 · strong-ASM 在 FP 富集 5×(OR=0.194, p=1.8e-28)· LOSO held-out mean −0.00004(p=0.125,方向相反)● L4重開須 materially new hypothesis + pre-decision audit;不影響 ISM 作為 characterization 工具的價值
★已測訊號主要與 germline-haplotype label 關聯HP1 vs HP2 association 較強;同 haplotype 內 somatic label 訊號弱(T3 AUC 0.65 < null 0.71)是方法層級限制,非 bug;不構成 unsupervised cluster 或 cellular subgroup truth
★全基因組 strong-ASM 極稀少0.34%(hypo 44% / hyper 56% 無方向偏好)★3稀有率直接限制任何應用方向的統計功效

「observed association」≠「能判別」——本研究最核心的張力

圖:具名 scope 觀察到 allele-associated methylation,但已測設計未形成 TP/FP 判別器
① observed association ✓ HP1 vs HP1-1 甲基明顯不同 HP1 HP1-1 scope-bound observed Δβ=−0.122 ② 用它「判別」TP/FP ✗ TP 與 FP 的 ASM 分數重疊 TP FP ASM 分數 → 重疊 → AUC=0.505 ≈ 隨機
左:在具名位點、既定 HP-axis 與分析 scope 中,觀察到 HP1 與 HP1-1 的 Δβ=−0.122;這是 association,不是 causal effect。右:在已測資料與 formulation 中,把全基因組位點按 ASM 分數排,TP 與 FP 的分布幾乎重疊(FP 甚至更常出現 strong-ASM)→ 未形成可用判別器。這就是「observed association ≠ 已證明能判別」。
結論 So-whatISM 目前有證據支持的定位 = 有明確 confound 邊界的 characterization 工具;已測設計未驗證它可作 discovery / filter。它用於描述:① observed 等位甲基差是否與 cis-compatible/copy-screen pattern 相符 ② 指定 HP/somatic label 是否與 read-distance variation 或 cluster × label 關聯 ③ 哪些位點統計受稀疏性限制。這些 association 不證明 cluster truth、cellular group、causal cis 或 tumor-acquired;新 filter 假說可經 pre-decision audit 另案重開。

誠實限制:normal-anchored cis heuristic 只在有 matched normal 的樣本完整(目前僅 HCC1395);依賴單一 pipeline → tier 上限 ★3。這是 project-specific combination;新穎性與比較優勢仍需具日期的 systematic literature review與 head-to-head benchmark。

§4 當時做法 vs 新做法 N-EVO

ISM 相對傳統/競品的三個關鍵差異。

面向當時做法(傳統/競品)新做法(ISM)差別 / 數據
甲基信號表示層modkit / pycoMeth:per-position 聚合,輸出每 CpG 的全 read 平均率per-read 表示,建 reads×CpG 矩陣,保留每 read 完整 pattern矩陣代表 between-molecule 資訊;聚合丟失 read 間共甲基化結構。數值 head-to-head {{待查}}
ASM confound screen直接比 HP1 vs HP2 率時可能混入 copy/germline-allelic confoundproject-specific normal-anchored cis heuristic 三路 + copy-pattern partition文獻:apparent ASM 有 82–92% 被 copy 數解釋(Martin-Trujillo);判準 |d_cis|>1.8|d_drift|+0.02 只標記部分 observed confound,不是完整校正
單倍型解析度NanoMethPhase 等:只到 HP1 vs HP2HP1 / HP1-1(somatic 子單倍型)/ HP2 / HP2-1HP1-1 依賴 longphase-S somatic tag,使 project-specific somatic-conditioned HP-axis 可行;不表示 confound 已完全控制
一句話差異聚合方法呈現「每個位點的平均甲基率」;ISM 另算「每對 read 之間的甲基距離」,並用 matched normal 的 project-specific heuristic 標記部分 copy-pattern confound;它不證明 confound 已扣除。

§5 難點拆解(最容易卡的點,都配圖) N-SVG

這幾個是新讀者最會卡的概念。難點 1-2 的圖在 §1 名詞卡內(Δ vs observed Δβ / HP1-1);以下是 cis heuristic 與 Cramér's V association gate。

難點 1:為什麼做 read×read 距離,而不是比 per-CpG 平均?

per-CpG 平均會把「read 之間的共甲基化結構」抹平。下圖:4 條 read × 5 個 CpG,HP1 兩條 read 彼此模式一致、HP2 兩條也一致但跟 HP1 相反。看每 CpG 平均(最下列)幾乎都 50%,看不出分群;看 read×read 距離才浮現兩亞群。

算式 worked example · NHD(read 間距離) NHD = 不一致數 / 共同有效 CpG 數(DistanceMatrix.cpp:23)。
read i = [1, 0, 1, 0, 1] · read j = [1, 0, 1, 1, 1] → 1 個不一致 / 5 共同 = 0.2
0 = 模式完全一致、1 = 完全相反。effective current CLI 需共同覆蓋 ≥ C_min=3 個 CpG,不足者依預設 SKIP 標記為無效對並排除。DistanceConfig internal initializer 5 及 legacy/experimental MAX_DIST=1.0 不是無參數生產行為。有效 CLI 無參數預設是 NHD;Config 雖以 BERNOULLI 初始化,會被參數解析器清空後回填 NHD。BERNOULLI(DistanceMatrix.cpp:254)只在明示指定時用 ML 甲基機率加權。
read 乘 CpG 矩陣轉成 read-read 距離 示意逐 read 甲基模式在 per-CpG 平均相同時仍可有不同的 read-level 距離結構。 read × CpG 矩陣 read×read 距離 聚類 → 2 群 C1C2C3C4C5 r1·HP1 r2·HP1 r3·HP2 r4·HP2 每CpG均 50%50%50%50%50% ↑ 平均看:全 50%,看不出分群 甲基化 未甲基 r1r2r3r4 r1r2r3r4 近(0) 遠(1) HP1 群HP2 群
(schematic)距離矩陣對角線=自己(近,深綠),r1-r2 / r3-r4 近,跨群遠(橘)→ 聚類分出兩群。per-CpG 平均(最下列全 50%)完全看不出——這就是 ISM 用 read×read 距離的理由。
難點 2:normal-anchored cis heuristic 怎麼標記 copy-screen-negative 候選

copy number 變化可能讓某 haplotype 份數與觀測平均甲基一起偏移。三路比較是 cis-compatible/copy-screen-negative heuristic,只標記與部分 observed copy pattern 相容或不相容的 candidate;不是完整 confound 校正或 validated causal test。

圖:normal 錨點三路比較
normal-anchored cis-test 三路比較 比較 normal-HP1、tumor-HP1 與 somatic-supporting tumor-HP1-1,用本專案 heuristic 標記 cis-compatible、copy-screen-negative candidate;不確認 cis truth。 normal-HP1 matched-normal observed baseline β≈基準 tumor-HP1 observed dosage-pattern candidate β↑ tumor-HP1-1 帶 somatic ALT β↓ d_copy(dosage screen) d_within(cis 候選) ▸ 判準 |d_cis| > 1.8×|d_drift| + 0.02。BRCA2:d_within=−0.023、d_copy=−0.11 → 與 observed copy pattern 相容 ▸ 例 chr17/TBC1D16:d_within=0.142, perm p=0.001 → cis-compatible candidate
(schematic)本專案用 normal-HP1 當錨點,把 normal→tumor-HP1 的 observed difference 記為 copy/dosage screen(d_copy),tumor-HP1→HP1-1 記為 cis-compatible 候選(d_within)。1.8×drift+0.02 是單樣本 heuristic,只 screen 部分 observed confound;不證明 copy/germline/drift 已移除、causal cis 或 tumor-acquired。僅 HCC1395 有 matched normal
難點 3:Cramér's V gate 與「487 association candidates」為何不是矛盾

稀疏 cluster × label 列聯表上 Cramér's V 不可信,ISM 用 Cochran 規則(至少 80% 格的期望值 ≥5)打 reliable 旗標;raw V 仍保留,summary projection 才歸 0,legacy gate 也未檢查此旗標。「summary CramérV=0 但 PERMANOVA 顯著」只表示 V 的 summary 投影不可用、而指定 label 的 read-distance association 在置換 null 下達門檻;須合讀 PERMDISP,不可升格為 unsupervised cluster existence/truth。

圖:稀疏表(gated)vs 密集表(可信)
Cochran reliability gate:稀疏表與密集表 稀疏列聯表的 Cramér's V 被 gate 為零;密集表才保留 V。PERMANOVA 顯著只是候選結構,仍需 dispersion 與設計檢查。 稀疏表:期望值 <5 的格超過 20% 2 1 1 1 HP1HP2 甲基未甲 → reliable = false → summary CramérV gated = 0 (PERMANOVA 仍可能顯著 → association candidate) 密集表:所有期望格 ≥ 5 30 8 6 25 HP1HP2 → reliable = true → CramérV 保留(可信)
(schematic)左稀疏 cluster × label 表的 Cramér's V 不可信 → gate 成 0;右密集表才保留 association strength。762 MISSED 中 487 個屬「gated=0 但指定 label 的 PERMANOVA 顯著」;這不是 cluster truth、cellular group 或因果。● 不可寫成 filter

§6 數據(每筆附 tier + 來源) N-CMP

本頁引用的關鍵數字。tier:● P 原檔對賬 ● S/caveat ● DEAD/NEGATIVE。

指標值tier來源
frozen baseline significance_summary.csv source 欄數199 欄● Prelease baseline ddd8909a source header;歷史輸出另有舊欄數,解析必須依欄名與 producer commit
PERMANOVA 置換次數(生產)999(p-floor 0.001)● Psrc/core/RegionProcessor.cpp + include/core/LabelTest.hpp
strong-ASM 全基因組稀有率0.34%(hypo 44%/hyper 56%)★3盤點線1 01_ism_method_spec_from_source.md L3
BRCA2 promoter 既定 HP-axis observed Δβ−0.122(與 observed copy pattern 相容;focal d_within=−0.023, p=0.024)● Pscope-bound association,非 causal effect;−0.054 為 buggy 砍半值,禁用
cis-compatible/copy-screen-negative candidate chr17/TBC1D16d_within=0.142, perm p=0.001★3project-specific heuristic;不證明 confound 已移除、causal cis 或 acquired
跨樣本 excess-over-null(3 癌種)6/6 正,mean 0.168;somatic private 0/38★3同上 L3
已測甲基→TP/FP filter formulationOR=0.194, p=1.8e-28;LOSO 100% circularity;AUC=0.505● L4 TESTED NEG.同上 L3;非 universal negative,新假說須另做 pre-decision audit
LOH-constrained phasing(NG=2 same-HP)93–99%,6/6 樣本;Wilcoxon p=0.0078B+ ★3同上 L3
O11 epipolymorphism(校正前→後)AUC 0.845 → 0.530(校正後=artifact)● NEG同上 L3