封存文章
本文源自 HiMCM 2024 的高功率運算題目。舊模型把 HPC、所有 data centres、 AI 與 cryptocurrency 的全球數據混成同一條 Gompertz curve,並輸出 0 MtCO₂e 及 1,484 Mt e-waste 等互相矛盾結果。本版本撤回該等點預測, 重建可審核的系統邊界與單位。
「高效能運算的足跡」沒有單一數字。訓練 AI model、運行氣候模擬、Bitcoin mining、 enterprise cloud 和 national supercomputer 的硬件、利用率、地點及服務功能都不同。 第一步不是揀一條 growth curve,而是定義:
- 評估對象是 facility、cluster、job 還是整個 sector?
- functional unit 是一次 job、每個 useful result、GPU-hour 還是年度服務?
- boundary 包括 operation、hardware manufacturing、building、network 和 end-of-life 到哪裏?
- energy、carbon、water、materials 和 local community impact 是否分開報告?
一、由工作量開始
對設備類別 和時間 :
其中 是 utilization, 是 utilization-to-power curve。 server idle power 通常不為零,不能簡單假設 energy 與 utilization 成正比。
若可直接量度 rack/PDU energy,應優先使用 meter:
Functional efficiency
只降低 kWh 未必代表服務更有效率。可報:
例如 converged simulations/kWh、validated samples/kWh, 並同時報告 result quality。FLOP/J 只描述硬件計算效率,不能保證演算法產生有用結果。
二、Facility energy 與 PUE
Power Usage Effectiveness:
所以:
PUE 包括 cooling、power distribution、lighting 等 overhead, 但不包含 embodied carbon,也不直接量度水或 computing productivity。 年度平均 PUE 會掩蓋季節與 part-load 效率;對 demand response 和 hourly carbon, 需要時間序列:
三、Operational carbon
Location-based emissions:
其中 單位可為 。 若 以 TWh:
因此:
這個簡潔換算可用來抓出舊程式的 744.2 TWh 卻顯示 0.0 MtCO₂e 的單位 bug。
Renewable accounting
若 已反映 grid mix,再乘 會 double count。 應分開:
- location-based grid emissions;
- market-based contractual accounting;
- onsite generation;
- temporal matching;
- curtailment/additionality。
對實際系統影響,可同時研究 average 與 marginal emissions;把年度 renewable certificate 等同每小時 zero-carbon supply 並不充分。
四、時間與地點
同一 job 在不同時段與地點運行:
carbon-aware scheduling 可以在 deadline、data locality 和 hardware constraint 下解:
限制每個 job 完成、capacity 不超限、deadline 達標。
但搬移 job 亦可能增加 network energy、延遲或需要 duplicate hardware; 不能只看一張靜態「哪個國家 carbon intensity 最低」的 bar chart。
五、水足跡
資料中心水要分至少兩層。
Onsite direct water
Water Usage Effectiveness:
蒸發冷卻的「consumption」與抽取後排回的「withdrawal」不同。 水質、季節、drift、blowdown 和 reclaimed water 都要記錄。
Electricity-supply water
發電技術的 water withdrawal/consumption 差異大。 舊模型用全局固定 乘所有 energy, 卻沒有說明是 onsite 或 power-sector water,容易 double count。
Scarcity weighting
一公升水在濕潤地區與 drought-stressed basin 的影響不同:
是有來源的 water-scarcity characterization factor。 因此單純「水量最少」與「水風險最低」未必同一方案。
六、Embodied emissions 與硬件
對 equipment cohort :
把 embodied impact 分配到 jobs 可按 service delivered:
但 allocation rule 必須公開。延長硬件壽命可降低 annualized manufacturing impact, 卻可能因舊硬件效率差而增加 operational energy;需要 break-even analysis。
Replacement decision
令舊、新設備每 unit work 的 energy 分別為 , 新設備 embodied carbon 為 。在 grid intensity 下, 碳回收所需工作量:
前提 。若預計剩餘工作量低於 ,提早更換未必減碳。
七、E-waste 不能用 spline 無限外推
全球 e-waste 包含家電、屏幕、電話及其他設備,不等於 HPC hardware。 要估算 computing equipment end-of-life,應用 cohort survival model:
是售出/安裝設備數, 是壽命分布。
Cubic spline 只保證 knots 之間平滑;超出最後資料點時可能劇烈發散。 舊模型一方面寫 2030 e-waste 70–90 Mt,實際程式又輸出 1,484.4 Mt, 已證明外推失控。這個數字不應保留。
e-waste 指標應包括:
- mass retired;
- reuse/refurbish fraction;
- certified recycling;
- recovered critical materials;
- hazardous treatment;
- data-destruction constraint。
八、Sector growth scenarios
Gompertz curve:
可以表達 saturating growth,但 不是由數學自然知道的 carrying capacity。 若只有少量歷史資料, 很難辨識,2030 結果會由 priors/assumptions 主導。
較透明的 decomposition:
其中:
- :service demand;
- :IT energy per unit service;
- :HPC、AI training/inference、enterprise、crypto 等明確 segment。
情景可分:
- demand growth;
- algorithmic efficiency;
- hardware efficiency;
- utilization;
- facility overhead;
- grid transition。
這讓讀者看見 rebound effect:
效率下降快過需求增長才會令總能耗下降。
九、不確定性
舊 Monte Carlo 對 growth、efficiency、renewable share、water intensity 指定獨立 normal distributions,再把負值 clip 至零。問題包括:
- 分布沒有資料來源;
- 參數相關性被忽略;
- clip 會在零形成 artificial mass;
- 只抽參數 uncertainty,沒有 model-form uncertainty;
- 稱 2.5–97.5 percentile 為「confidence interval」。
可建立分層情景:
由來源支持 bounded/lognormal/beta distributions,保留相關性, 並將結果稱為 scenario uncertainty interval 或 posterior predictive interval。
Global sensitivity
使用 Sobol 或 variance decomposition:
它比逐一 ±10% 更能處理 interaction。
十、Local impacts 與公平
Global carbon 低不代表 local impact 低。facility-level assessment 應有:
- grid interconnection 和 backup generators;
- local air pollutants;
- noise;
- land use;
- water basin 和 competing demand;
- heat rejection;
- tax/employment;
- electricity/water infrastructure cost allocation;
- community consultation。
舊圖用自行設定的 0–10 severity scores 畫 local effects 和 socioeconomic impacts, 只能作 criteria inventory,不能稱為量度。每一項需要可觀測 indicator、 affected population 和 evidence source。
十一、改善方案如何量化
1. Software/algorithm
比較同等 output quality 下:
包括 lower precision、better solver、early stopping、model reuse 和 right-sizing。 不能以較差結果換取低能耗後仍稱同一服務。
2. Hardware utilization
- consolidate idle servers;
- power management;
- accelerator matching;
- batch scheduling;
- reuse before replacement。
3. Cooling
評估 PUE、WUE、climate、water scarcity、heat reuse 和 refrigerants; 沒有一種 cooling 在所有地區同時最省水、最省電。
4. Electricity
- hourly carbon-aware scheduling;
- additional clean generation;
- transmission/storage constraint;
- backup generator emissions;
- credible location- and market-based reporting。
5. Circularity
- modular repair;
- component reuse;
- procurement longevity requirements;
- take-back;
- audited recycling;
- material traceability。
十二、指標 dashboard
| 層次 | 指標 |
|---|---|
| Job | kWh/job、kgCO₂e/job、quality、runtime |
| Cluster | utilization、idle energy、throughput/kWh |
| Facility | IT MWh、PUE、peak MW、onsite WUE |
| Supply | hourly CI、source-water intensity |
| Hardware | embodied kgCO₂e、age、reuse/recycle |
| Community | water scarcity、noise、backup emissions |
所有指標需同時提供 denominator;總量和 intensity 應一起看,避免 efficiency improvement 掩蓋 absolute growth。
十三、模型驗證
- IT meter 與 facility meter 的 energy balance;
- PUE denominator 一致;
- TWh→kWh→MtCO₂e 單位測試;
- location/market accounting 不 double count;
- onsite 和 source water 分開;
- hardware cohort mass balance;
- sector data 不與 global e-waste 重複;
- scenario 參數有 provenance;
- historical backcast/holdout;
- 結果能重算,不用手填 bar scores。
十四、舊 2030 輸出的處理
舊程式報:
- energy:744.2 TWh;
- water:1,339,539 million litres;
- carbon:0.0 MtCO₂e;
- e-waste:1,484.4 Mt。
後兩者已顯示單位/spline failure,前兩者亦由任意 Gompertz 和固定 water intensity 產生。 所以不能保留「450–750 TWh」、「減排 30–60%」或「水增 25–40%」 作研究發現。它們只可列為舊情景假設的範圍。
結論
計算環境足跡最重要的不是建立最複雜的綜合分數,而是保持邊界、功能單位和能源帳透明:
- 由 service demand 與 measured IT energy 開始;
- 用 PUE 連接 facility overhead;
- 以 hourly/location-specific electricity 計 operational carbon;
- 分開 onsite 與 electricity-supply water;
- 用 hardware cohorts 計 embodied impact 和 e-waste;
- 以情景呈現 demand、efficiency 和 grid 的不確定性;
- 同時報 absolute total 與 per-service intensity。
修正後,這篇文章不再聲稱已預測 2030 的全球 HPC 足跡, 而是提供一套可以由實際 meter、workload 和 hardware inventory 填入的核算框架。
參考資料
- COMAP, Inc. (2024). HiMCM 2024 Problem B and accompanying resources.
- The Green Grid. PUE: A Comprehensive Examination of the Metric.
- ISO. (2006). ISO 14040 and ISO 14044: Life-cycle assessment principles and requirements.
- International Telecommunication Union & UNITAR. (2024). The Global E-waste Monitor 2024.
- Mytton, D. (2021). Data centre water consumption. npj Clean Water, 4, 11.
- Shehabi, A., Smith, S. J., Sartor, D. A., et al. (2016). United States Data Center Energy Usage Report. Lawrence Berkeley National Laboratory.