A Sutro-group pitch, written Tue 22 Sep 2026 (week 39), from an idea dictated Mon 21 Sep 13:11 — and independently named in the room the same evening: "It's like a thermally aware compiler." Prior art checked, own measurements re-verified, and the experiment that would kill it.
What you would be re-inventing (say this first, before someone says it for you)
- Thermal-aware compilation is patented, twice. US 9,639,359 "Thermal-aware compiler for parallel instruction execution" and US 8,972,957 "Thermal-aware source code compilation". Binding instructions to the coolest functional unit has a name in the literature (TempIB); there is an ACM Computing Surveys review from 2012.
- "Temperature is an architectural variable" is Skadron et al., ISCA 2003 (HotSpot) — 23 years old, and their own retrospective says its main use became thermally-aware task scheduling.
- "The power budget, not the transistor count, is what you schedule against" is dark silicon, ISCA 2011.
- Frequency as a first-class compiler variable is already shipping in ML: Zeus, Perseus (per-microbatch DVFS against pipeline slack), DynamoLLM, PALS.
- The datacenter half is solved at production scale: Facebook Dynamo (ISCA 2016), Google Thunderbolt (OSDI 2020) — 25%+ power oversubscription in production, without breaking the compiler abstraction.
- 🛑 And the one that lands closest to your ET-SoC-1 result: ALUPower measured a single GPU kernel swinging 155 W → 257 W purely on operand values, all-zero vectors at the bottom. Your "38 W on zeros vs 67 W on random" is a reproduction on new silicon, not a discovery. Say so in the first minute and you keep the room; let someone else say it and you lose the talk.
What actually survives
- Data-dependent energy as a compiler cost model. ALUPower measured it as a phenomenon; nobody compiles against it. A scheduler that knows
zeros ≪ constants ≪ randomcan choose tilings, layouts and encodings to minimise bit flips. That lane is open. - Doing it on silicon with no DVFS loop. Every prior system schedules against a controller. The ET-SoC-1 has none, so the physics lands on the energy axis undistorted. You have the only bench like this in the room.
- The framing — your own line from yesterday's post brief: you cannot remove the leak, you can only choose which axis it comes out on. Regulate power and it lands on frequency; pin frequency and it lands entirely on joules. That is sharper than "heat-aware compiler," which is 2003.
Your three measurements, re-verified against the raw files
| Measurement | Status |
|---|---|
| A100 8192³: fp16 sustains 1,216.2 MHz vs the 1,410 MHz boost (−13.7%, +15.9% wall clock); int8 1,274.3 MHz; fp32 and the b1 AND+POPC kernel hold exactly 1,410.0 MHz in every sample at the same 54–60 °C | ✅ confirmed in clocks_sm_MHz |
| At the measured clock the kernel is at 93.7% of scaled peak, not 80.8% of datasheet peak — two thirds of "cuBLAS inefficiency" is the clock, not the code | ✅ |
| Two A100s, same part number: 57.7 W vs 37.5 W idle (55%), ~1% speed difference | ⚠️ true, but the boards differ in power cap (400 W vs 320 W), driver and VBIOS — not a clean part-to-part comparison, and the 1.19% speed delta is inside your own 4.2% noise |
| ET-SoC-1: clock pinned at 600 MHz, fp32 TensorFMA 546 cycles/op regardless of data, 38.0 W zeros → 47.6 W constants → 64.5 W uniform → 67.3 W random (~1.8×) | ✅ two rounds agree within 1.1 W |
| "The power cap caused the 1,216 MHz" | 🛑 inference, not measurement. All 90 recorded throttle-reason masks decode to NONE because every one is an idle snapshot. Your hottest sample anywhere is 66 °C against an 85 °C target, so no thermal event ever fired. One 100 ms-sampled nvmlDeviceGetCurrentClocksThrottleReasons run closes this — and it is the same one-line fix yesterday's post brief flagged |
The pitch, as it should be spoken
A compiler that knows where the fan is.
We built a wall: the chip designer guarantees the chip stays up, the compiler pretends physics doesn't exist. Both sides then lie to each other. Three lies I measured on my own hardware this month.
One: my A100 runs an fp16 tensor-core GEMM at 1,216 MHz, not the 1,410 every cycle-count model assumes — 15.9% of wall clock, gone. Two: on the same board, in the same run, at the same temperature, fp32 and a binary AND+POPC kernel lose zero megahertz. Temperature isn't the variable; power is. Three: on the ET-SoC-1 — a chip with no DVFS loop at all — fp32 matmul takes 546 cycles per op regardless of the data, and burns 38 W on zeros, 67 W on random normal. Same instructions, same speed, 1.8× the energy, purely from how many transistors flip.
So the statement isn't "abstractions leak." It's: you cannot remove the leak, you can only choose which axis it comes out on.
Why now: thermal-aware compilation has been patented since 2015 and nobody uses it, because the payoff was milliseconds on a laptop. That changed — training clusters now swing hundreds of megawatts and the grid is the binding constraint — and the compiler's search space is no longer human-sized. We can afford to search placement × schedule × datatype × bit-flip count, which is exactly the chunk that was "too hard" in 2003.
Three questions it must answer: (1) what is the cost model — joules per op as a function of operand values, not opcode? (2) what does it control — placement, clock, encoding, or the shape of the power waveform in time? (3) what does it optimise — time, energy, peak power, or dP/dt?
The smallest experiment that kills it, one afternoon, on hardware you already reach: pin the A100's SM clock across seven steps on the existing Modal harness, rerun the 8192³ fp16 GEMM, and log throttle reasons inside the 100 ms sampler. Clean line in 1/f + SwPowerCap ⇒ there is a budget to schedule against. HwThermalSlowdown ⇒ the idea is about cooling, not code. In parallel on the ET-SoC-1: same matmul, operands permuted to minimise Hamming distance. If joules move and cycles don't, a compiler has a free lunch it is not taking.
The strongest objection, and the answer
Rupert Wu already gave it out loud at #32: "you're only shaving on milliseconds." Naveen Rao's version would be blunter — the hyperscalers solved this at the layer where the money is, without breaking any abstraction.
Concede the time axis entirely. Yes: on speed this is milliseconds and the datacenter people own it. The claim is on the energy axis, where nobody is scheduling: two identically-named A100s differ 55% in idle power, >99% of a short MNIST run's joules are the idle floor, and on the ET-SoC-1 the same instruction sequence at the same cycle count costs 1.8× depending on operand bits. Runtime DVFS cannot touch that — by the time the runtime sees the workload, the bits are already chosen. Only a compiler picks the layout, the tiling and the encoding that decide how many transistors flip.
"An instruction takes N cycles" leaks 14%. "An instruction costs N picojoules" leaks 80%.
That sentence is the talk, and it is the one number in your data nobody else in that room has measured.