Gyroscope report

A compiler that knows where the fan is

A Sutro-group pitch, written Tue 22 Sep 2026 (week 39), from an idea dictated Mon 21 Sep 13:11 — and independently named in the room the same evening: "It's like a thermally aware compiler." Prior art checked, own measurements re-verified, and the experiment that would kill it.

What you would be re-inventing (say this first, before someone says it for you)

What actually survives

  1. Data-dependent energy as a compiler cost model. ALUPower measured it as a phenomenon; nobody compiles against it. A scheduler that knows zeros ≪ constants ≪ random can choose tilings, layouts and encodings to minimise bit flips. That lane is open.
  2. Doing it on silicon with no DVFS loop. Every prior system schedules against a controller. The ET-SoC-1 has none, so the physics lands on the energy axis undistorted. You have the only bench like this in the room.
  3. The framing — your own line from yesterday's post brief: you cannot remove the leak, you can only choose which axis it comes out on. Regulate power and it lands on frequency; pin frequency and it lands entirely on joules. That is sharper than "heat-aware compiler," which is 2003.

Your three measurements, re-verified against the raw files

Measurement Status
A100 8192³: fp16 sustains 1,216.2 MHz vs the 1,410 MHz boost (−13.7%, +15.9% wall clock); int8 1,274.3 MHz; fp32 and the b1 AND+POPC kernel hold exactly 1,410.0 MHz in every sample at the same 54–60 °C ✅ confirmed in clocks_sm_MHz
At the measured clock the kernel is at 93.7% of scaled peak, not 80.8% of datasheet peak — two thirds of "cuBLAS inefficiency" is the clock, not the code ✅
Two A100s, same part number: 57.7 W vs 37.5 W idle (55%), ~1% speed difference ⚠️ true, but the boards differ in power cap (400 W vs 320 W), driver and VBIOS — not a clean part-to-part comparison, and the 1.19% speed delta is inside your own 4.2% noise
ET-SoC-1: clock pinned at 600 MHz, fp32 TensorFMA 546 cycles/op regardless of data, 38.0 W zeros → 47.6 W constants → 64.5 W uniform → 67.3 W random (~1.8×) ✅ two rounds agree within 1.1 W
"The power cap caused the 1,216 MHz" 🛑 inference, not measurement. All 90 recorded throttle-reason masks decode to NONE because every one is an idle snapshot. Your hottest sample anywhere is 66 °C against an 85 °C target, so no thermal event ever fired. One 100 ms-sampled nvmlDeviceGetCurrentClocksThrottleReasons run closes this — and it is the same one-line fix yesterday's post brief flagged

The pitch, as it should be spoken

A compiler that knows where the fan is.

We built a wall: the chip designer guarantees the chip stays up, the compiler pretends physics doesn't exist. Both sides then lie to each other. Three lies I measured on my own hardware this month.

One: my A100 runs an fp16 tensor-core GEMM at 1,216 MHz, not the 1,410 every cycle-count model assumes — 15.9% of wall clock, gone. Two: on the same board, in the same run, at the same temperature, fp32 and a binary AND+POPC kernel lose zero megahertz. Temperature isn't the variable; power is. Three: on the ET-SoC-1 — a chip with no DVFS loop at all — fp32 matmul takes 546 cycles per op regardless of the data, and burns 38 W on zeros, 67 W on random normal. Same instructions, same speed, 1.8× the energy, purely from how many transistors flip.

So the statement isn't "abstractions leak." It's: you cannot remove the leak, you can only choose which axis it comes out on.

Why now: thermal-aware compilation has been patented since 2015 and nobody uses it, because the payoff was milliseconds on a laptop. That changed — training clusters now swing hundreds of megawatts and the grid is the binding constraint — and the compiler's search space is no longer human-sized. We can afford to search placement × schedule × datatype × bit-flip count, which is exactly the chunk that was "too hard" in 2003.

Three questions it must answer: (1) what is the cost model — joules per op as a function of operand values, not opcode? (2) what does it control — placement, clock, encoding, or the shape of the power waveform in time? (3) what does it optimise — time, energy, peak power, or dP/dt?

The smallest experiment that kills it, one afternoon, on hardware you already reach: pin the A100's SM clock across seven steps on the existing Modal harness, rerun the 8192³ fp16 GEMM, and log throttle reasons inside the 100 ms sampler. Clean line in 1/f + SwPowerCap ⇒ there is a budget to schedule against. HwThermalSlowdown ⇒ the idea is about cooling, not code. In parallel on the ET-SoC-1: same matmul, operands permuted to minimise Hamming distance. If joules move and cycles don't, a compiler has a free lunch it is not taking.

The strongest objection, and the answer

Rupert Wu already gave it out loud at #32: "you're only shaving on milliseconds." Naveen Rao's version would be blunter — the hyperscalers solved this at the layer where the money is, without breaking any abstraction.

Concede the time axis entirely. Yes: on speed this is milliseconds and the datacenter people own it. The claim is on the energy axis, where nobody is scheduling: two identically-named A100s differ 55% in idle power, >99% of a short MNIST run's joules are the idle floor, and on the ET-SoC-1 the same instruction sequence at the same cycle count costs 1.8× depending on operand bits. Runtime DVFS cannot touch that — by the time the runtime sees the workload, the bits are already chosen. Only a compiler picks the layout, the tiling and the encoding that decide how many transistors flip.

"An instruction takes N cycles" leaks 14%. "An instruction costs N picojoules" leaks 80%.

That sentence is the talk, and it is the one number in your data nobody else in that room has measured.