intermediate
The Benchmark That Only Wins for Ten Seconds
Read the counters before the options. Nothing here is labelled with the answer.
The report
Our new vectorized encoder benchmarks 35% faster than the old one. In production it is about 4% faster, and on the busiest nodes it is slower. The benchmark runs the same input on the same instance type. We have run it dozens of times and it is consistent.
CountersSIMULATED
| benchmark duration | 8 s per run | Each benchmark run completes in a few seconds. |
| core frequency, first 5 s of benchmark | ≈ 4.6 GHz | The core runs near its maximum clock at the start of a run. |
| core frequency, after 60 s sustained load | ≈ 2.9 GHz | Under continuous load the clock settles substantially lower. |
| frequency during wide-vector sections | ≈ 2.6 GHz | The clock is lower again while wide vector instructions are executing. |
| instructions retired | new version 0.62× the old | The new version retires substantially fewer instructions for the same work. |
| package power | at the configured limit under sustained load | The processor is drawing as much power as it is permitted to draw. |
| benchmark: cores active | 1 of 32 | The benchmark exercises a single core. |
| production: cores active | 28–32 of 32 | Production runs the workload on nearly every core at once. |
What is the hardware doing?