Doc machine code budgets - #121
Merged
Merged
Conversation
…ode budget Doc-only. A LogCompilation probe of the jitAudit kernels (Microsoft OpenJDK 25.0.4, x86-64, 8-lane double) contradicts three claims the repo currently makes. The measurements, reading `stub_offset - insts_offset` off each c2 <nmethod>: Object.<init> 1 bytecode -> 208 bytes machine code doublearrays.clamp! 180 -> 1728 (9.6x, 69% of 2500) doublearrays.meanAndVarianceTwoPass 248 -> 1688 (6.8x, 68%) vecxt.all.clamp! (export forwarder) 11 -> 1696 (154x, 68%) ~200 bytes of fixed overhead plus 7-10x the bytecode for vectorised code with a masked tail. At that ratio `InlineSmallCode` (2500, machine code) is reached at roughly 260-300 bytecodes — *below* `FreqInlineSize` (325, bytecode). For a SIMD kernel the machine-code limit binds first. HotPath.java claimed check C2 meant "C2 will inline it into its callers once it is hot". It does not: that is a necessary condition, not a sufficient one, and the doc now says so and points at D2 for the compiled size. Thin.java asserted a 35-bytecode budget as the thing that makes the public API zero-cost at a cold call site. `vecxt.all.clamp!` satisfies it by a factor of three while being 1696 bytes of machine code. Forwarders are where the ratio is most extreme, precisely because the body they forward to gets pulled in — and the `vecxt.all` forwarders are excluded from the baseline by `primaryAnnotated`, so their compiled size is unmeasured twice over. The blog gets a new subsection under the threshold table, "The limit that is not measured in bytecodes", because every threshold in that table is a bytecode count and this one is not. Also records that `FreqInlineSize` and `InlineSmallCode` are `C2 pd product` — platform-dependent — so two of the four budgets the page relies on are properties of the machine rather than of HotSpot. And that C1/C2/C3 read bytecode and therefore cannot see any of this. `doublearrays.variance(mode)` gets the specific consequence written down: its `@AllocFree` zero depends on C2 inlining `meanAndVarianceTwoPass`, which is at 68% of the limit. If that crosses, the pair escapes and the symptom is a D1 failure complaining about allocation rather than about inlining. Two things deliberately not claimed. The ratio is one workload on one CPU at one lane width, so the direction is established and the crossover point approximate. And whether HotSpot's MaxTrivialSize/MaxInlineSize fast paths let a small callee bypass the InlineSmallCode veto is unverified — both docs say so rather than guessing, because it decides whether the forwarder finding is a curiosity or a hazard. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tch @AllocFree Three things, one story: the LogCompilation-parsing checks are not being built, so the annotations and docs that pointed at them as future work are corrected, and the check that does cover the same ground gets its population widened. == D2 and D5 will not be built == The probe settled feasibility and then undermined the case. <intrinsic> is emitted and self-describing, and a Mill test fork produces a complete log — so D2 is buildable. But the vocabulary is JDK-internal with no compatibility contract (the plan's guessed inline_fail strings were wrong in five places), and the drift is indistinguishable from the regression: a renamed intrinsic id shrinks the observed set, which is exactly what losing vectorisation looks like, with no invariant to violate and no cross-check available. The line that fell out of it is which API a check depends on. D1 uses ThreadMXBean, D3 uses VectorSpecies, D4 uses arithmetic, eaOffTest uses a product-grade -XX: flag — all public and contractual. D2 and D5 would have depended on the compiler's diagnostic output instead. That is the dividing line, not the amount of work. Recorded in jitAudit/package.mill rather than left as an absence, because an unexplained gap in a checklist reads as an oversight. The forward references written in the previous commit — HotPath.java's "check D2's business", the variance comment's "see check D2", the blog's pointer — now say the limit is documented and unenforced, and name what covers it indirectly. == eaOffTest is what replaces D2, and it is switched off == Reading it properly while answering "do these modules earn their keep": D1 with EA enabled cannot distinguish "the SIMD intrinsics were applied" from "they were not, and EA scalarised the software-path objects instead" — both read as zero. The EA-off scope can, because on the software path with EA off the objects reach the heap. That is D2's question answered from an observable consequence rather than from XML, it is already written, and its CI step is commented out. Its scaladoc now says so. Nothing removed. == D1 widened from 20 to 29 kernels == Every @AllocFree method now has a test behind it. The nine added: floatarrays clamp!, +=(Array[Float]), *=(Array[Float]), *=(Float) intarrays minSIMD, maxSIMD, +=(Array[Int]), -=(Array[Int]), -=(Int) An annotation with no test is an assertion nobody has checked, which is how intarrays.dot carried @AllocFree while allocating a dead array per call. The class doc now states that coverage is every annotated method and that the two are kept in step by hand — and drops the brittle counts that have needed correcting twice. Also notes why the kernel lists cannot be factored into a shared collection: assertAllocFree must stay inline so each call site gets a monomorphic measurement loop, and driving them from a List[() => Unit] would reintroduce the megamorphic Function0.apply() dispatch that stops C2 inlining through to the Vector API calls. The duplication is load-bearing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Bytecode audit (Tier 1) —
|
| threshold | value | provenance |
|---|---|---|
MaxTrivialSize |
6 | discovered |
MaxInlineSize |
35 | discovered |
FreqInlineSize |
325 | discovered |
MaxInlineLevel |
15 | discovered |
InlineSmallCode |
2500 | discovered |
NodeCountInliningCutoff |
18000 | assumed |
HugeMethodLimit |
8000 | assumed |
-XX:HugeMethodLimit= was rejected on the command line: a develop flag compiled out of this product build, so 8000 is taken from the HotSpot source and cannot be confirmed against the running JVM.
| metric | now | baseline | delta |
|---|---|---|---|
| cheatsheet methods | 97 | — | — |
| total bytes | 42067 | 51103 | -17.7% |
| distinct library ops | 198 | 178 | +11.2% |
| bytes per op | 212.5 | 287.1 | -26.0% |
| severity | check | at | method | detail |
|---|---|---|---|---|
| WARN | C1 | cheatsheet.scala:118 |
CheatsheetTest$.matrixRangeSlicing |
5643 bytes, past 68% of HugeMethodLimit=8000 (assumed) |
Annotated methods (164)
| method | annotations | bytes | budget used | loop | at |
|---|---|---|---|---|---|
vecxt.intarrays$.$plus |
@Thin |
24 | 69% of 35 | no | intarrays.scala:546 |
vecxt.intarrays$.$minus |
@Thin |
24 | 69% of 35 | no | intarrays.scala:399 |
vecxt.NDArrayFloatOps$.compareGeneral |
@HotPath |
195 | 60% of 325 | yes | ndarrayFloatOps.scala:67 |
vecxt.NDArrayFloatOps$.binaryOpGeneral |
@HotPath |
195 | 60% of 325 | yes | ndarrayFloatOps.scala:20 |
vecxt.NDArrayIntOps$.compareGeneral |
@HotPath |
186 | 57% of 325 | yes | ndarrayIntOps.scala:67 |
vecxt.NDArrayIntOps$.binaryOpGeneral |
@HotPath |
186 | 57% of 325 | yes | ndarrayIntOps.scala:19 |
vecxt.NDArrayDoubleOps$.compareGeneral |
@HotPath |
186 | 57% of 325 | yes | ndarrayDoubleOps.scala:74 |
vecxt.NDArrayDoubleOps$.binaryOpGeneral |
@HotPath |
186 | 57% of 325 | yes | ndarrayDoubleOps.scala:24 |
vecxt.doublearrays$.clamp$bang |
@AllocFree @HotPath |
180 | 55% of 325 | yes | doublearrays.scala:872 |
vecxt.matrix$Layout.linearIndex |
@Thin |
19 | 54% of 35 | no | matrix.scala:48 |
vecxt.floatarrays$.clamp$bang |
@AllocFree @HotPath |
172 | 53% of 325 | yes | floatarrays.scala:417 |
vecxt.NDArrayFloatOps$.compareScalarGeneral |
@HotPath |
165 | 51% of 325 | yes | ndarrayFloatOps.scala:94 |
vecxt.NDArrayFloatOps$.binaryOpInPlaceGeneral |
@HotPath |
162 | 50% of 325 | yes | ndarrayFloatOps.scala:119 |
vecxt.NDArrayDoubleOps$.compareScalarGeneral |
@HotPath |
157 | 48% of 325 | yes | ndarrayDoubleOps.scala:102 |
vecxt.NDArrayIntOps$.compareScalarGeneral |
@HotPath |
156 | 48% of 325 | yes | ndarrayIntOps.scala:94 |
vecxt.NDArrayIntOps$.binaryOpInPlaceGeneral |
@HotPath |
153 | 47% of 325 | yes | ndarrayIntOps.scala:119 |
vecxt.NDArrayDoubleOps$.binaryOpInPlaceGeneral |
@HotPath |
153 | 47% of 325 | yes | ndarrayDoubleOps.scala:131 |
vecxt.NDArrayIntOps$.unaryOpGeneral |
@HotPath |
152 | 47% of 325 | yes | ndarrayIntOps.scala:42 |
vecxt.NDArrayDoubleOps$.unaryOpGeneral |
@HotPath |
152 | 47% of 325 | yes | ndarrayDoubleOps.scala:48 |
vecxt.intarrays$.$minus |
@Thin |
16 | 46% of 35 | no | intarrays.scala:519 |
vecxt.doublearrays$.fillLinspace |
@AllocFree @HotPath |
133 | 41% of 325 | yes | doublearrays.scala:34 |
vecxt.intarrays$.increments |
@HotPath |
121 | 37% of 325 | yes | intarrays.scala:219 |
vecxt.ndarray$.mkNDArray |
@Thin |
13 | 37% of 35 | no | ndarray.scala:188 |
vecxt.doublearrays$.increments |
@HotPath |
110 | 34% of 325 | yes | doublearrays.scala:397 |
vecxt.intarrays$.dot |
@AllocFree @HotPath |
108 | 33% of 325 | yes | intarrays.scala:372 |
vecxt.doublearrays$.unary_$minus |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:172 |
vecxt.doublearrays$.tanh |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:365 |
vecxt.doublearrays$.tan |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:354 |
vecxt.doublearrays$.sqrt |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:317 |
vecxt.doublearrays$.sinh |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:339 |
vecxt.doublearrays$.sin |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:328 |
vecxt.doublearrays$.log1p |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:306 |
vecxt.doublearrays$.log10 |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:295 |
vecxt.doublearrays$.log |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:284 |
vecxt.doublearrays$.expm1 |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:273 |
vecxt.doublearrays$.exp |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:262 |
vecxt.doublearrays$.cosh |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:251 |
vecxt.doublearrays$.cos |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:240 |
vecxt.doublearrays$.cbrt |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:229 |
vecxt.doublearrays$.atan |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:218 |
vecxt.doublearrays$.asin |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:207 |
vecxt.doublearrays$.acos |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:196 |
vecxt.doublearrays$.abs |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:184 |
vecxt.intarrays$.$less |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:48 |
vecxt.intarrays$.$less$eq |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:52 |
vecxt.intarrays$.$greater |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:56 |
vecxt.intarrays$.$greater$eq |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:60 |
vecxt.intarrays$.$eq$colon$eq |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:40 |
vecxt.intarrays$.$bang$colon$eq |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:44 |
vecxt.floatarrays$.unary_$minus |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:114 |
vecxt.floatarrays$.tanh |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:303 |
vecxt.floatarrays$.tan |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:292 |
vecxt.floatarrays$.sqrt |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:259 |
vecxt.floatarrays$.sinh |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:281 |
vecxt.floatarrays$.sin |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:270 |
vecxt.floatarrays$.log1p |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:248 |
vecxt.floatarrays$.log10 |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:237 |
vecxt.floatarrays$.log |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:226 |
vecxt.floatarrays$.expm1 |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:215 |
vecxt.floatarrays$.exp |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:204 |
vecxt.floatarrays$.cosh |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:193 |
vecxt.floatarrays$.cos |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:182 |
vecxt.floatarrays$.cbrt |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:171 |
vecxt.floatarrays$.atan |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:160 |
vecxt.floatarrays$.asin |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:149 |
vecxt.floatarrays$.acos |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:138 |
vecxt.floatarrays$.abs |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:126 |
vecxt.intarrays$.mean |
@Thin |
11 | 31% of 35 | no | intarrays.scala:287 |
vecxt.doublearrays$.tanh$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.tan$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.sqrt$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.sinh$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.sin$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.log1p$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.log10$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.log$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.expm1$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.exp$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.cosh$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.cos$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.cbrt$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.atan$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.asin$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.acos$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.abs$bang |
@AllocFree @HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.$minus$bang |
@AllocFree @HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.floatarrays$.increments |
@HotPath |
97 | 30% of 325 | yes | floatarrays.scala:652 |
vecxt.doublearrays$.$times$times$bang |
@HotPath |
95 | 29% of 325 | yes | doublearrays.scala:372 |
vecxt.intarrays$.$less |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:139 |
vecxt.intarrays$.$less$eq |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:143 |
vecxt.intarrays$.$greater |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:147 |
vecxt.intarrays$.$greater$eq |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:151 |
vecxt.intarrays$.$eq$colon$eq |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:131 |
vecxt.intarrays$.$bang$colon$eq |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:135 |
vecxt.floatarrays$.tanh$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.tan$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.sqrt$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.sinh$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.sin$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.log1p$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.log10$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.log$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.expm1$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.exp$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.cosh$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.cos$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.cbrt$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.atan$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.asin$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.acos$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.abs$bang |
@AllocFree @HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.floatarrays$.$minus$bang |
@AllocFree @HotPath |
93 | 29% of 325 | yes | floatarrays.scala:94 |
vecxt.intarrays$.variance |
@Thin |
10 | 29% of 35 | no | intarrays.scala:304 |
vecxt.intarrays$.std |
@Thin |
10 | 29% of 35 | no | intarrays.scala:361 |
vecxt.doublearrays$.variance |
@AllocFree @Thin |
10 | 29% of 35 | no | doublearrays.scala:541 |
vecxt.doublearrays$.$minus$eq |
@AllocFree @HotPath |
92 | 28% of 325 | yes | doublearrays.scala:1063 |
vecxt.doublearrays$.$plus$eq |
@AllocFree @HotPath |
90 | 28% of 325 | yes | doublearrays.scala:998 |
vecxt.doublearrays$.$times$eq |
@AllocFree @HotPath |
88 | 27% of 325 | yes | doublearrays.scala:1135 |
vecxt.doublearrays$.productSIMD |
@AllocFree @HotPath |
86 | 26% of 325 | yes | doublearrays.scala:700 |
vecxt.floatarrays$.$times$times$bang |
@HotPath |
85 | 26% of 325 | yes | floatarrays.scala:314 |
vecxt.doublearrays$.sumSIMD |
@AllocFree @HotPath |
85 | 26% of 325 | yes | doublearrays.scala:677 |
vecxt.doublearrays$.fma$bang |
@AllocFree @HotPath |
85 | 26% of 325 | yes | doublearrays.scala:1038 |
vecxt.intarrays$.$plus$eq |
@AllocFree @HotPath |
84 | 26% of 325 | yes | intarrays.scala:555 |
vecxt.intarrays$.$minus$eq |
@AllocFree @HotPath |
84 | 26% of 325 | yes | intarrays.scala:527 |
vecxt.floatarrays$.$times$eq |
@AllocFree @HotPath |
84 | 26% of 325 | yes | floatarrays.scala:874 |
vecxt.floatarrays$.$plus$eq |
@AllocFree @HotPath |
84 | 26% of 325 | yes | floatarrays.scala:772 |
vecxt.floatarrays$.$minus$eq |
@AllocFree @HotPath |
84 | 26% of 325 | yes | floatarrays.scala:812 |
vecxt.ndarrayOps.expandDims |
@Thin |
9 | 26% of 35 | no | ndarrayOps.scala |
vecxt.intarrays.stdDev |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.meanAndVariance |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.lte |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.lte |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.lt |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.lt |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.gte |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.gte |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.gt |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.gt |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays$.variance |
@Thin |
9 | 26% of 35 | no | intarrays.scala:291 |
vecxt.intarrays$.stdDev |
@Thin |
9 | 26% of 35 | no | intarrays.scala:364 |
vecxt.intarrays$.std |
@Thin |
9 | 26% of 35 | no | intarrays.scala:357 |
vecxt.intarrays$.meanAndVariance |
@Thin |
9 | 26% of 35 | no | intarrays.scala:308 |
vecxt.doublearrays.meanAndVariance |
@Thin |
9 | 26% of 35 | no | doublearrays.scala |
vecxt.doublearrays$.meanAndVariance |
@Thin |
9 | 26% of 35 | no | doublearrays.scala:555 |
vecxt.intarrays$.minSIMD |
@AllocFree @HotPath |
82 | 25% of 325 | yes | intarrays.scala:575 |
vecxt.intarrays$.maxSIMD |
@AllocFree @HotPath |
82 | 25% of 325 | yes | intarrays.scala:595 |
vecxt.floatarrays$.productSIMD |
@AllocFree @HotPath |
82 | 25% of 325 | yes | floatarrays.scala:537 |
vecxt.intarrays$.sumSIMD |
@AllocFree @HotPath |
81 | 25% of 325 | yes | intarrays.scala:265 |
vecxt.floatarrays$.sumSIMD |
@AllocFree @HotPath |
81 | 25% of 325 | yes | floatarrays.scala:516 |
vecxt.floatarrays$.fma$bang |
@AllocFree @HotPath |
80 | 25% of 325 | yes | floatarrays.scala:340 |
vecxt.floatarrays$.$times$eq |
@AllocFree @HotPath |
78 | 24% of 325 | yes | floatarrays.scala:926 |
vecxt.ndarray.shapeArray |
@Thin |
8 | 23% of 35 | no | ndarray.scala |
vecxt.matrix$Matrix.rows |
@Thin |
8 | 23% of 35 | no | matrix.scala:129 |
vecxt.matrix$Matrix.rowStride |
@Thin |
8 | 23% of 35 | no | matrix.scala:135 |
vecxt.matrix$Matrix.offset |
@Thin |
8 | 23% of 35 | no | matrix.scala:141 |
vecxt.matrix$Matrix.numel |
@Thin |
8 | 23% of 35 | no | matrix.scala:144 |
vecxt.matrix$Matrix.isDenseRowMajor |
@Thin |
8 | 23% of 35 | no | matrix.scala:150 |
vecxt.matrix$Matrix.isDenseColMajor |
@Thin |
8 | 23% of 35 | no | matrix.scala:147 |
vecxt.matrix$Matrix.hasSimpleContiguousMemoryLayout |
@Thin |
8 | 23% of 35 | no | matrix.scala:162 |
vecxt.matrix$Matrix.cols |
@Thin |
8 | 23% of 35 | no | matrix.scala:132 |
vecxt.matrix$Matrix.colStride |
@Thin |
8 | 23% of 35 | no | matrix.scala:138 |
vecxt.intarrays$.countsToIdx |
@HotPath |
70 | 22% of 325 | yes | intarrays.scala:244 |
vecxt.intarrays$.$minus$eq |
@AllocFree @HotPath |
67 | 21% of 325 | yes | intarrays.scala:408 |
vecxt.floatarrays$.$plus$eq |
@AllocFree @HotPath |
24 | 7% of 325 | no | floatarrays.scala:731 |
Method sizes
| band | methods |
|---|---|
| <= 6 (trivial, always inlined) | 1411 |
| 7-35 (inlinable cold) | 3569 |
| 36-325 (inlinable when hot) | 702 |
| 326-8000 (not inlined) | 90 |
| > 8000 (NEVER JIT COMPILED) | 0 |
| bytes | method | module | at |
|---|---|---|---|
| 5643 | CheatsheetTest$.matrixRangeSlicing |
experiments | cheatsheet.scala:118 |
| 3815 | CheatsheetTest$.matrixReverseSlicing |
experiments | cheatsheet.scala:125 |
| 3801 | CheatsheetTest$.ndArrayInt |
experiments | cheatsheet.scala:416 |
| 3779 | CheatsheetTest$.ndArrayBoolean |
experiments | cheatsheet.scala:430 |
| 3061 | CheatsheetTest$.ndArrayFloat |
experiments | cheatsheet.scala:395 |
| 3015 | CheatsheetTest$.ndArrayFloatReductions |
experiments | cheatsheet.scala:405 |
| 2811 | CheatsheetTest$.matrixCreation |
experiments | cheatsheet.scala:72 |
| 2597 | CheatsheetTest$.matrixOps |
experiments | cheatsheet.scala:167 |
| 2500 | vecxt_re.Tower.show |
vecxt_re | Tower.scala:63 |
| 1594 | vecxt.ndarrayOps$.apply |
vecxt | ndarrayOps.scala:452 |
| 1460 | CheatsheetTest$.indexingAndSlicing |
experiments | cheatsheet.scala:98 |
| 1435 | vecxt.JvmDoubleMatrix$.$plus$eq |
vecxt | doublematrix.scala:315 |
| 1420 | vecxt.JvmFloatMatrix$.floatmatrixAddScalarInPlace |
vecxt | floatmatrix.scala:442 |
| 1420 | vecxt.JvmFloatMatrix$.floatmatrixSubScalarInPlace |
vecxt | floatmatrix.scala:514 |
| 1308 | CheatsheetTest$.arrayManipulation |
experiments | cheatsheet.scala:305 |
| 1122 | vecxt.Svd$.pinv |
vecxt | svd.scala:42 |
| 1040 | vecxt.JvmFloatMatrix$.floatmatrixSubVector |
vecxt | floatmatrix.scala:336 |
| 1028 | CheatsheetTest$.matrixFloat |
experiments | cheatsheet.scala:445 |
| 984 | vecxt.JvmDoubleMatrix$.$plus$eq |
vecxt | doublematrix.scala:227 |
| 969 | vecxt_re.Scenarr$.combine |
vecxt_re | scenarr.scala:160 |
| 960 | vecxt.JvmFloatMatrix$.floatmatrixAddVectorInPlace |
vecxt | floatmatrix.scala:248 |
| 960 | vecxt.JvmFloatMatrix$.floatmatrixSubVectorInPlace |
vecxt | floatmatrix.scala:360 |
| 923 | CheatsheetTest$.matrixInt |
experiments | cheatsheet.scala:461 |
| 908 | vecxt.Svd$.svd |
vecxt | svd.scala:135 |
| 896 | vecxt_re.NegativeBinomial$.volweightedMle |
vecxt_re | NegativeBinomial.scala:281 |
Proposed baseline
{
"jdkMajor": 25,
"c9": { "totalBytes": 42067, "distinctOps": 198 },
"annotated": {
"vecxt.NDArrayDoubleOps$.binaryOpGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
"vecxt.NDArrayDoubleOps$.binaryOpInPlaceGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)V": 153,
"vecxt.NDArrayDoubleOps$.compareGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
"vecxt.NDArrayDoubleOps$.compareScalarGeneral(Lvecxt/ndarray$NDArray;DLscala/Function2;)Lvecxt/ndarray$NDArray;": 157,
"vecxt.NDArrayDoubleOps$.unaryOpGeneral(Lvecxt/ndarray$NDArray;Lscala/Function1;)Lvecxt/ndarray$NDArray;": 152,
"vecxt.NDArrayFloatOps$.binaryOpGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 195,
"vecxt.NDArrayFloatOps$.binaryOpInPlaceGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)V": 162,
"vecxt.NDArrayFloatOps$.compareGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 195,
"vecxt.NDArrayFloatOps$.compareScalarGeneral(Lvecxt/ndarray$NDArray;FLscala/Function2;)Lvecxt/ndarray$NDArray;": 165,
"vecxt.NDArrayIntOps$.binaryOpGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
"vecxt.NDArrayIntOps$.binaryOpInPlaceGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)V": 153,
"vecxt.NDArrayIntOps$.compareGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
"vecxt.NDArrayIntOps$.compareScalarGeneral(Lvecxt/ndarray$NDArray;ILscala/Function2;)Lvecxt/ndarray$NDArray;": 156,
"vecxt.NDArrayIntOps$.unaryOpGeneral(Lvecxt/ndarray$NDArray;Lscala/Function1;)Lvecxt/ndarray$NDArray;": 152,
"vecxt.doublearrays$.$minus$bang([D)V": 98,
"vecxt.doublearrays$.$minus$eq([DD)V": 92,
"vecxt.doublearrays$.$plus$eq([DD)V": 90,
"vecxt.doublearrays$.$times$eq([D[D)V": 88,
"vecxt.doublearrays$.$times$times$bang([DD)V": 95,
"vecxt.doublearrays$.abs$bang([D)V": 98,
"vecxt.doublearrays$.abs([D)[D": 108,
"vecxt.doublearrays$.acos$bang([D)V": 98,
"vecxt.doublearrays$.acos([D)[D": 108,
"vecxt.doublearrays$.asin$bang([D)V": 98,
"vecxt.doublearrays$.asin([D)[D": 108,
"vecxt.doublearrays$.atan$bang([D)V": 98,
"vecxt.doublearrays$.atan([D)[D": 108,
"vecxt.doublearrays$.cbrt$bang([D)V": 98,
"vecxt.doublearrays$.cbrt([D)[D": 108,
"vecxt.doublearrays$.clamp$bang([DDD)V": 180,
"vecxt.doublearrays$.cos$bang([D)V": 98,
"vecxt.doublearrays$.cos([D)[D": 108,
"vecxt.doublearrays$.cosh$bang([D)V": 98,
"vecxt.doublearrays$.cosh([D)[D": 108,
"vecxt.doublearrays$.exp$bang([D)V": 98,
"vecxt.doublearrays$.exp([D)[D": 108,
"vecxt.doublearrays$.expm1$bang([D)V": 98,
"vecxt.doublearrays$.expm1([D)[D": 108,
"vecxt.doublearrays$.fillLinspace([DDD)V": 133,
"vecxt.doublearrays$.fma$bang([DDD)V": 85,
"vecxt.doublearrays$.increments([D)[D": 110,
"vecxt.doublearrays$.log$bang([D)V": 98,
"vecxt.doublearrays$.log([D)[D": 108,
"vecxt.doublearrays$.log10$bang([D)V": 98,
"vecxt.doublearrays$.log10([D)[D": 108,
"vecxt.doublearrays$.log1p$bang([D)V": 98,
"vecxt.doublearrays$.log1p([D)[D": 108,
"vecxt.doublearrays$.meanAndVariance([D)Lvecxt/MeanAndVariance;": 9,
"vecxt.doublearrays$.productSIMD([D)D": 86,
"vecxt.doublearrays$.sin$bang([D)V": 98,
"vecxt.doublearrays$.sin([D)[D": 108,
"vecxt.doublearrays$.sinh$bang([D)V": 98,
"vecxt.doublearrays$.sinh([D)[D": 108,
"vecxt.doublearrays$.sqrt$bang([D)V": 98,
"vecxt.doublearrays$.sqrt([D)[D": 108,
"vecxt.doublearrays$.sumSIMD([D)D": 85,
"vecxt.doublearrays$.tan$bang([D)V": 98,
"vecxt.doublearrays$.tan([D)[D": 108,
"vecxt.doublearrays$.tanh$bang([D)V": 98,
"vecxt.doublearrays$.tanh([D)[D": 108,
"vecxt.doublearrays$.unary_$minus([D)[D": 108,
"vecxt.doublearrays$.variance([DLvecxt/VarianceMode;)D": 10,
"vecxt.doublearrays.meanAndVariance([DLvecxt/VarianceMode;)Lvecxt/MeanAndVariance;": 9,
"vecxt.floatarrays$.$minus$bang([F)V": 93,
"vecxt.floatarrays$.$minus$eq([FF)V": 84,
"vecxt.floatarrays$.$plus$eq([FF)V": 84,
"vecxt.floatarrays$.$plus$eq([F[F)V": 24,
"vecxt.floatarrays$.$times$eq([FF)V": 78,
"vecxt.floatarrays$.$times$eq([F[F)V": 84,
"vecxt.floatarrays$.$times$times$bang([FF)V": 85,
"vecxt.floatarrays$.abs$bang([F)V": 93,
"vecxt.floatarrays$.abs([F)[F": 104,
"vecxt.floatarrays$.acos$bang([F)V": 93,
"vecxt.floatarrays$.acos([F)[F": 104,
"vecxt.floatarrays$.asin$bang([F)V": 93,
"vecxt.floatarrays$.asin([F)[F": 104,
"vecxt.floatarrays$.atan$bang([F)V": 93,
"vecxt.floatarrays$.atan([F)[F": 104,
"vecxt.floatarrays$.cbrt$bang([F)V": 93,
"vecxt.floatarrays$.cbrt([F)[F": 104,
"vecxt.floatarrays$.clamp$bang([FFF)V": 172,
"vecxt.floatarrays$.cos$bang([F)V": 93,
"vecxt.floatarrays$.cos([F)[F": 104,
"vecxt.floatarrays$.cosh$bang([F)V": 93,
"vecxt.floatarrays$.cosh([F)[F": 104,
"vecxt.floatarrays$.exp$bang([F)V": 93,
"vecxt.floatarrays$.exp([F)[F": 104,
"vecxt.floatarrays$.expm1$bang([F)V": 93,
"vecxt.floatarrays$.expm1([F)[F": 104,
"vecxt.floatarrays$.fma$bang([FFF)V": 80,
"vecxt.floatarrays$.increments([F)[F": 97,
"vecxt.floatarrays$.log$bang([F)V": 93,
"vecxt.floatarrays$.log([F)[F": 104,
"vecxt.floatarrays$.log10$bang([F)V": 93,
"vecxt.floatarrays$.log10([F)[F": 104,
"vecxt.floatarrays$.log1p$bang([F)V": 93,
"vecxt.floatarrays$.log1p([F)[F": 104,
"vecxt.floatarrays$.productSIMD([F)F": 82,
"vecxt.floatarrays$.sin$bang([F)V": 93,
"vecxt.floatarrays$.sin([F)[F": 104,
"vecxt.floatarrays$.sinh$bang([F)V": 93,
"vecxt.floatarrays$.sinh([F)[F": 104,
"vecxt.floatarrays$.sqrt$bang([F)V": 93,
"vecxt.floatarrays$.sqrt([F)[F": 104,
"vecxt.floatarrays$.sumSIMD([F)F": 81,
"vecxt.floatarrays$.tan$bang([F)V": 93,
"vecxt.floatarrays$.tan([F)[F": 104,
"vecxt.floatarrays$.tanh$bang([F)V": 93,
"vecxt.floatarrays$.tanh([F)[F": 104,
"vecxt.floatarrays$.unary_$minus([F)[F": 104,
"vecxt.intarrays$.$bang$colon$eq([II)[Z": 94,
"vecxt.intarrays$.$bang$colon$eq([I[I)[Z": 106,
"vecxt.intarrays$.$eq$colon$eq([II)[Z": 94,
"vecxt.intarrays$.$eq$colon$eq([I[I)[Z": 106,
"vecxt.intarrays$.$greater$eq([II)[Z": 94,
"vecxt.intarrays$.$greater$eq([I[I)[Z": 106,
"vecxt.intarrays$.$greater([II)[Z": 94,
"vecxt.intarrays$.$greater([I[I)[Z": 106,
"vecxt.intarrays$.$less$eq([II)[Z": 94,
"vecxt.intarrays$.$less$eq([I[I)[Z": 106,
"vecxt.intarrays$.$less([II)[Z": 94,
"vecxt.intarrays$.$less([I[I)[Z": 106,
"vecxt.intarrays$.$minus$eq([II)V": 67,
"vecxt.intarrays$.$minus$eq([I[I)V": 84,
"vecxt.intarrays$.$minus([II)[I": 16,
"vecxt.intarrays$.$minus([I[I)[I": 24,
"vecxt.intarrays$.$plus$eq([I[I)V": 84,
"vecxt.intarrays$.$plus([I[I)[I": 24,
"vecxt.intarrays$.countsToIdx([I)[I": 70,
"vecxt.intarrays$.dot([I[I)I": 108,
"vecxt.intarrays$.increments([I)[I": 121,
"vecxt.intarrays$.maxSIMD([I)I": 82,
"vecxt.intarrays$.mean([I)D": 11,
"vecxt.intarrays$.meanAndVariance([I)Lvecxt/MeanAndVariance;": 9,
"vecxt.intarrays$.minSIMD([I)I": 82,
"vecxt.intarrays$.std([I)D": 9,
"vecxt.intarrays$.std([ILvecxt/VarianceMode;)D": 10,
"vecxt.intarrays$.stdDev([I)D": 9,
"vecxt.intarrays$.sumSIMD([I)I": 81,
"vecxt.intarrays$.variance([I)D": 9,
"vecxt.intarrays$.variance([ILvecxt/VarianceMode;)D": 10,
"vecxt.intarrays.gt([II)[Z": 9,
"vecxt.intarrays.gt([I[I)[Z": 9,
"vecxt.intarrays.gte([II)[Z": 9,
"vecxt.intarrays.gte([I[I)[Z": 9,
"vecxt.intarrays.lt([II)[Z": 9,
"vecxt.intarrays.lt([I[I)[Z": 9,
"vecxt.intarrays.lte([II)[Z": 9,
"vecxt.intarrays.lte([I[I)[Z": 9,
"vecxt.intarrays.meanAndVariance([ILvecxt/VarianceMode;)Lvecxt/MeanAndVariance;": 9,
"vecxt.intarrays.stdDev([ILvecxt/VarianceMode;)D": 9,
"vecxt.matrix$Layout.linearIndex(II)I": 19,
"vecxt.matrix$Matrix.colStride()I": 8,
"vecxt.matrix$Matrix.cols()I": 8,
"vecxt.matrix$Matrix.hasSimpleContiguousMemoryLayout()Z": 8,
"vecxt.matrix$Matrix.isDenseColMajor()Z": 8,
"vecxt.matrix$Matrix.isDenseRowMajor()Z": 8,
"vecxt.matrix$Matrix.numel()I": 8,
"vecxt.matrix$Matrix.offset()I": 8,
"vecxt.matrix$Matrix.rowStride()I": 8,
"vecxt.matrix$Matrix.rows()I": 8,
"vecxt.ndarray$.mkNDArray(Ljava/lang/Object;[I[II)Lvecxt/ndarray$NDArray;": 13,
"vecxt.ndarray.shapeArray(Lvecxt/ndarray$NDArray;)[I": 8,
"vecxt.ndarrayOps.expandDims(Lvecxt/ndarray$NDArray;I)Lvecxt/ndarray$NDArray;": 9
}
}
…y with D1 The reason it sat commented out since #110 turns out to be structural rather than incidental: it had no canary. Every assertion in it reads "still zero with EA off", which is precisely what a run with EA still *on* produces — so the scope passed whether or not -XX:-DoEscapeAnalysis reached the JVM. That is not a check, and switching it off was the right call at the time. The canary was available for free and nobody had noticed. `variance(mode)` is the one kernel in D1Suite whose zero comes from escape analysis rather than from intrinsification: it reads one field out of the MeanAndVariance that meanAndVarianceTwoPass returns and discards the other, so the object is dead and EA removes it. That object is an ordinary final class, not a Vector, so nothing intrinsifies it away — with EA off it must reach the heap. Asserting that it *does* allocate proves the flag took effect, and a failure there says every other assertion in the scope is passing for the wrong reason. That is also why it is the canary rather than a 29th kernel assertion: asserting "≤ 8 bytes/op with EA off" for an EA-dependent kernel would be asserting the opposite of what the flag does. The distinction the whole scope rests on is intrinsification-eliminated (survives EA-off) versus EA-eliminated (does not), and variance is the only member of the second category. Coverage 14 -> 28 kernels plus the canary, so every @AllocFree method is now asserted in both scopes. The drift is worth noting as a hazard in itself: this suite sat at fourteen while D1Suite grew to twenty-nine, and the lists cannot be factored into a shared collection because the assertion helpers must stay inline for each call site to get a monomorphic measurement loop. Docs corrected in two places where I had overstated this scope's reach. It relies on EA being what rescues a software-path kernel, and that is not always so — D6's canary is a software-path kernel that allocates with EA *enabled*, so D1Suite catches that one unaided. What this scope adds is the narrower set of fallbacks whose objects happen not to escape. Cheap and worth having; not the full substitute for confirming intrinsification that "what replaces D2" implied. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.