De-inline the cumsum pair and the BLAS forwarders - #124
Merged
Conversation
Six methods per file, in doublearrays and its floatarrays mirror. None of them meets any of the three conditions the inlining policy keeps `inline` for: no closure parameter, no VectorOperators constant to hold, and a concrete element type at every point they touch an element. cumsum! @hotpath @AllocFree cumsum @thin dot @thin norm @thin - @thin -= @thin (the Array overload; see below) `cumsum!` is the interesting one. A prefix sum carries a dependency — `vec(i)` needs the value just written to `vec(i - 1)` — so it cannot be vectorised and there is no Vector API in the body at all. That makes it the only `@AllocFree` kernel in the library whose zero is unconditional rather than contingent on C2 applying intrinsics: there is no DoubleVector whose failure to scalarise could put bytes on the heap. It is worth measuring precisely because it is the control case for every other kernel in D1 — if it ever fails, the harness is wrong and not the kernel. Added to D1Suite and D1EAOffSuite for both element types, which brings both suites to 31 tests and keeps every @AllocFree method covered in both. `dot`, `norm` and the array-argument `-=` are `@Thin` rather than `@HotPath`, and the distinction is the annotation's own wording: `@HotPath` describes "code that runs once per element", and the per-element loop in these is inside netlib's ddot/dnrm2/daxpy, not in the method. What the method does is name the operation and dispatch, which is `@Thin`'s definition. C3's no-backward-branch assertion holds for the same reason and would fail if anyone open-coded the loop back in, which is the right outcome. None of the three is `@AllocFree`. `blas` is `JavaBLAS.getInstance`, so these call into a third-party pure-Java implementation whose allocation behaviour is not visible from here and has never been measured. Both annotations that turned out to be false — `**!` and `intarrays.dot` — were applied by inspection, so inspection is not the standard being used here. Worth a comment where it lands: `-=(Array)` is a BLAS forwarder while `-=(Double)` is a hand-written Vector API loop carrying @hotpath @AllocFree. Same name, different implementations, so different annotations, and not an inconsistency for someone to tidy up later. Scope note: four of these were named directly; `norm` and `-` were added because they sit inside the same block and are the same two shapes, and leaving them `inline` would half-convert it. `add`, `+` and `+=` in the same file are the same shapes again but are left alone — `add` forwards to `+`, so it cannot become a clean forwarder until `+` is emitted, and that ordering is a batch of its own. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Bytecode audit (Tier 1) —
|
| threshold | value | provenance |
|---|---|---|
MaxTrivialSize |
6 | discovered |
MaxInlineSize |
35 | discovered |
FreqInlineSize |
325 | discovered |
MaxInlineLevel |
15 | discovered |
InlineSmallCode |
2500 | discovered |
NodeCountInliningCutoff |
18000 | assumed |
HugeMethodLimit |
8000 | assumed |
-XX:HugeMethodLimit= was rejected on the command line: a develop flag compiled out of this product build, so 8000 is taken from the HotSpot source and cannot be confirmed against the running JVM.
| metric | now | baseline | delta |
|---|---|---|---|
| cheatsheet methods | 97 | — | — |
| total bytes | 41872 | 51103 | -18.1% |
| distinct library ops | 206 | 178 | +15.7% |
| bytes per op | 203.3 | 287.1 | -29.2% |
| severity | check | at | method | detail |
|---|---|---|---|---|
| WARN | C1 | cheatsheet.scala:118 |
CheatsheetTest$.matrixRangeSlicing |
5643 bytes, past 68% of HugeMethodLimit=8000 (assumed) |
Annotated methods (176)
| method | annotations | bytes | budget used | loop | at |
|---|---|---|---|---|---|
vecxt.floatarrays$.$minus$eq |
@Thin |
26 | 74% of 35 | no | floatarrays.scala:735 |
vecxt.doublearrays$.$minus$eq |
@Thin |
26 | 74% of 35 | no | doublearrays.scala:999 |
vecxt.intarrays$.$plus |
@Thin |
24 | 69% of 35 | no | intarrays.scala:546 |
vecxt.intarrays$.$minus |
@Thin |
24 | 69% of 35 | no | intarrays.scala:399 |
vecxt.floatarrays$.$minus |
@Thin |
24 | 69% of 35 | no | floatarrays.scala:724 |
vecxt.doublearrays$.$minus |
@Thin |
24 | 69% of 35 | no | doublearrays.scala:987 |
vecxt.floatarrays$.dot |
@Thin |
23 | 66% of 35 | no | floatarrays.scala:650 |
vecxt.doublearrays$.dot |
@Thin |
23 | 66% of 35 | no | doublearrays.scala:978 |
vecxt.NDArrayFloatOps$.compareGeneral |
@HotPath |
195 | 60% of 325 | yes | ndarrayFloatOps.scala:67 |
vecxt.NDArrayFloatOps$.binaryOpGeneral |
@HotPath |
195 | 60% of 325 | yes | ndarrayFloatOps.scala:20 |
vecxt.NDArrayIntOps$.compareGeneral |
@HotPath |
186 | 57% of 325 | yes | ndarrayIntOps.scala:67 |
vecxt.NDArrayIntOps$.binaryOpGeneral |
@HotPath |
186 | 57% of 325 | yes | ndarrayIntOps.scala:19 |
vecxt.NDArrayDoubleOps$.compareGeneral |
@HotPath |
186 | 57% of 325 | yes | ndarrayDoubleOps.scala:74 |
vecxt.NDArrayDoubleOps$.binaryOpGeneral |
@HotPath |
186 | 57% of 325 | yes | ndarrayDoubleOps.scala:24 |
vecxt.doublearrays$.clamp$bang |
@AllocFree @HotPath |
180 | 55% of 325 | yes | doublearrays.scala:872 |
vecxt.matrix$Layout.linearIndex |
@Thin |
19 | 54% of 35 | no | matrix.scala:48 |
vecxt.floatarrays$.clamp$bang |
@AllocFree @HotPath |
172 | 53% of 325 | yes | floatarrays.scala:418 |
vecxt.NDArrayFloatOps$.compareScalarGeneral |
@HotPath |
165 | 51% of 325 | yes | ndarrayFloatOps.scala:94 |
vecxt.NDArrayFloatOps$.binaryOpInPlaceGeneral |
@HotPath |
162 | 50% of 325 | yes | ndarrayFloatOps.scala:119 |
vecxt.NDArrayDoubleOps$.compareScalarGeneral |
@HotPath |
157 | 48% of 325 | yes | ndarrayDoubleOps.scala:102 |
vecxt.NDArrayIntOps$.compareScalarGeneral |
@HotPath |
156 | 48% of 325 | yes | ndarrayIntOps.scala:94 |
vecxt.NDArrayIntOps$.binaryOpInPlaceGeneral |
@HotPath |
153 | 47% of 325 | yes | ndarrayIntOps.scala:119 |
vecxt.NDArrayDoubleOps$.binaryOpInPlaceGeneral |
@HotPath |
153 | 47% of 325 | yes | ndarrayDoubleOps.scala:131 |
vecxt.NDArrayIntOps$.unaryOpGeneral |
@HotPath |
152 | 47% of 325 | yes | ndarrayIntOps.scala:42 |
vecxt.NDArrayDoubleOps$.unaryOpGeneral |
@HotPath |
152 | 47% of 325 | yes | ndarrayDoubleOps.scala:48 |
vecxt.intarrays$.$minus |
@Thin |
16 | 46% of 35 | no | intarrays.scala:519 |
vecxt.floatarrays$.cumsum |
@Thin |
15 | 43% of 35 | no | floatarrays.scala:694 |
vecxt.doublearrays$.cumsum |
@Thin |
15 | 43% of 35 | no | doublearrays.scala:959 |
vecxt.doublearrays$.fillLinspace |
@AllocFree @HotPath |
133 | 41% of 325 | yes | doublearrays.scala:34 |
vecxt.intarrays$.increments |
@HotPath |
121 | 37% of 325 | yes | intarrays.scala:219 |
vecxt.ndarray$.mkNDArray |
@Thin |
13 | 37% of 35 | no | ndarray.scala:188 |
vecxt.floatarrays$.norm |
@Thin |
13 | 37% of 35 | no | floatarrays.scala:655 |
vecxt.doublearrays$.norm |
@Thin |
13 | 37% of 35 | no | doublearrays.scala:983 |
vecxt.doublearrays$.increments |
@HotPath |
110 | 34% of 325 | yes | doublearrays.scala:397 |
vecxt.intarrays$.dot |
@AllocFree @HotPath |
108 | 33% of 325 | yes | intarrays.scala:372 |
vecxt.doublearrays$.unary_$minus |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:172 |
vecxt.doublearrays$.tanh |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:365 |
vecxt.doublearrays$.tan |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:354 |
vecxt.doublearrays$.sqrt |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:317 |
vecxt.doublearrays$.sinh |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:339 |
vecxt.doublearrays$.sin |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:328 |
vecxt.doublearrays$.log1p |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:306 |
vecxt.doublearrays$.log10 |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:295 |
vecxt.doublearrays$.log |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:284 |
vecxt.doublearrays$.expm1 |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:273 |
vecxt.doublearrays$.exp |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:262 |
vecxt.doublearrays$.cosh |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:251 |
vecxt.doublearrays$.cos |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:240 |
vecxt.doublearrays$.cbrt |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:229 |
vecxt.doublearrays$.atan |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:218 |
vecxt.doublearrays$.asin |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:207 |
vecxt.doublearrays$.acos |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:196 |
vecxt.doublearrays$.abs |
@HotPath |
108 | 33% of 325 | yes | doublearrays.scala:184 |
vecxt.intarrays$.$less |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:48 |
vecxt.intarrays$.$less$eq |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:52 |
vecxt.intarrays$.$greater |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:56 |
vecxt.intarrays$.$greater$eq |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:60 |
vecxt.intarrays$.$eq$colon$eq |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:40 |
vecxt.intarrays$.$bang$colon$eq |
@HotPath |
106 | 33% of 325 | yes | intarrays.scala:44 |
vecxt.floatarrays$.unary_$minus |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:115 |
vecxt.floatarrays$.tanh |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:304 |
vecxt.floatarrays$.tan |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:293 |
vecxt.floatarrays$.sqrt |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:260 |
vecxt.floatarrays$.sinh |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:282 |
vecxt.floatarrays$.sin |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:271 |
vecxt.floatarrays$.log1p |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:249 |
vecxt.floatarrays$.log10 |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:238 |
vecxt.floatarrays$.log |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:227 |
vecxt.floatarrays$.expm1 |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:216 |
vecxt.floatarrays$.exp |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:205 |
vecxt.floatarrays$.cosh |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:194 |
vecxt.floatarrays$.cos |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:183 |
vecxt.floatarrays$.cbrt |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:172 |
vecxt.floatarrays$.atan |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:161 |
vecxt.floatarrays$.asin |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:150 |
vecxt.floatarrays$.acos |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:139 |
vecxt.floatarrays$.abs |
@HotPath |
104 | 32% of 325 | yes | floatarrays.scala:127 |
vecxt.intarrays$.mean |
@Thin |
11 | 31% of 35 | no | intarrays.scala:287 |
vecxt.doublearrays$.tanh$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.tan$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.sqrt$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.sinh$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.sin$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.log1p$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.log10$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.log$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.expm1$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.exp$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.cosh$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.cos$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.cbrt$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.atan$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.asin$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.acos$bang |
@HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.abs$bang |
@AllocFree @HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.doublearrays$.$minus$bang |
@AllocFree @HotPath |
98 | 30% of 325 | yes | doublearrays.scala:151 |
vecxt.floatarrays$.increments |
@HotPath |
97 | 30% of 325 | yes | floatarrays.scala:659 |
vecxt.doublearrays$.$times$times$bang |
@HotPath |
95 | 29% of 325 | yes | doublearrays.scala:372 |
vecxt.intarrays$.$less |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:139 |
vecxt.intarrays$.$less$eq |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:143 |
vecxt.intarrays$.$greater |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:147 |
vecxt.intarrays$.$greater$eq |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:151 |
vecxt.intarrays$.$eq$colon$eq |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:131 |
vecxt.intarrays$.$bang$colon$eq |
@HotPath |
94 | 29% of 325 | yes | intarrays.scala:135 |
vecxt.floatarrays$.tanh$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.tan$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.sqrt$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.sinh$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.sin$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.log1p$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.log10$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.log$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.expm1$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.exp$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.cosh$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.cos$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.cbrt$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.atan$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.asin$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.acos$bang |
@HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.abs$bang |
@AllocFree @HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.floatarrays$.$minus$bang |
@AllocFree @HotPath |
93 | 29% of 325 | yes | floatarrays.scala:95 |
vecxt.intarrays$.variance |
@Thin |
10 | 29% of 35 | no | intarrays.scala:304 |
vecxt.intarrays$.std |
@Thin |
10 | 29% of 35 | no | intarrays.scala:361 |
vecxt.doublearrays$.variance |
@AllocFree @Thin |
10 | 29% of 35 | no | doublearrays.scala:541 |
vecxt.doublearrays$.$minus$eq |
@AllocFree @HotPath |
92 | 28% of 325 | yes | doublearrays.scala:1093 |
vecxt.doublearrays$.$plus$eq |
@AllocFree @HotPath |
90 | 28% of 325 | yes | doublearrays.scala:1028 |
vecxt.doublearrays$.$times$eq |
@AllocFree @HotPath |
88 | 27% of 325 | yes | doublearrays.scala:1165 |
vecxt.doublearrays$.productSIMD |
@AllocFree @HotPath |
86 | 26% of 325 | yes | doublearrays.scala:700 |
vecxt.floatarrays$.$times$times$bang |
@HotPath |
85 | 26% of 325 | yes | floatarrays.scala:315 |
vecxt.doublearrays$.sumSIMD |
@AllocFree @HotPath |
85 | 26% of 325 | yes | doublearrays.scala:677 |
vecxt.doublearrays$.fma$bang |
@AllocFree @HotPath |
85 | 26% of 325 | yes | doublearrays.scala:1068 |
vecxt.intarrays$.$plus$eq |
@AllocFree @HotPath |
84 | 26% of 325 | yes | intarrays.scala:555 |
vecxt.intarrays$.$minus$eq |
@AllocFree @HotPath |
84 | 26% of 325 | yes | intarrays.scala:527 |
vecxt.floatarrays$.$times$eq |
@AllocFree @HotPath |
84 | 26% of 325 | yes | floatarrays.scala:892 |
vecxt.floatarrays$.$plus$eq |
@AllocFree @HotPath |
84 | 26% of 325 | yes | floatarrays.scala:790 |
vecxt.floatarrays$.$minus$eq |
@AllocFree @HotPath |
84 | 26% of 325 | yes | floatarrays.scala:830 |
vecxt.ndarrayOps.expandDims |
@Thin |
9 | 26% of 35 | no | ndarrayOps.scala |
vecxt.intarrays.stdDev |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.meanAndVariance |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.lte |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.lte |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.lt |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.lt |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.gte |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.gte |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.gt |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays.gt |
@Thin |
9 | 26% of 35 | no | intarrays.scala |
vecxt.intarrays$.variance |
@Thin |
9 | 26% of 35 | no | intarrays.scala:291 |
vecxt.intarrays$.stdDev |
@Thin |
9 | 26% of 35 | no | intarrays.scala:364 |
vecxt.intarrays$.std |
@Thin |
9 | 26% of 35 | no | intarrays.scala:357 |
vecxt.intarrays$.meanAndVariance |
@Thin |
9 | 26% of 35 | no | intarrays.scala:308 |
vecxt.doublearrays.meanAndVariance |
@Thin |
9 | 26% of 35 | no | doublearrays.scala |
vecxt.doublearrays$.meanAndVariance |
@Thin |
9 | 26% of 35 | no | doublearrays.scala:555 |
vecxt.intarrays$.minSIMD |
@AllocFree @HotPath |
82 | 25% of 325 | yes | intarrays.scala:575 |
vecxt.intarrays$.maxSIMD |
@AllocFree @HotPath |
82 | 25% of 325 | yes | intarrays.scala:595 |
vecxt.floatarrays$.productSIMD |
@AllocFree @HotPath |
82 | 25% of 325 | yes | floatarrays.scala:538 |
vecxt.intarrays$.sumSIMD |
@AllocFree @HotPath |
81 | 25% of 325 | yes | intarrays.scala:265 |
vecxt.floatarrays$.sumSIMD |
@AllocFree @HotPath |
81 | 25% of 325 | yes | floatarrays.scala:517 |
vecxt.floatarrays$.fma$bang |
@AllocFree @HotPath |
80 | 25% of 325 | yes | floatarrays.scala:341 |
vecxt.floatarrays$.$times$eq |
@AllocFree @HotPath |
78 | 24% of 325 | yes | floatarrays.scala:944 |
vecxt.ndarray.shapeArray |
@Thin |
8 | 23% of 35 | no | ndarray.scala |
vecxt.matrix$Matrix.rows |
@Thin |
8 | 23% of 35 | no | matrix.scala:129 |
vecxt.matrix$Matrix.rowStride |
@Thin |
8 | 23% of 35 | no | matrix.scala:135 |
vecxt.matrix$Matrix.offset |
@Thin |
8 | 23% of 35 | no | matrix.scala:141 |
vecxt.matrix$Matrix.numel |
@Thin |
8 | 23% of 35 | no | matrix.scala:144 |
vecxt.matrix$Matrix.isDenseRowMajor |
@Thin |
8 | 23% of 35 | no | matrix.scala:150 |
vecxt.matrix$Matrix.isDenseColMajor |
@Thin |
8 | 23% of 35 | no | matrix.scala:147 |
vecxt.matrix$Matrix.hasSimpleContiguousMemoryLayout |
@Thin |
8 | 23% of 35 | no | matrix.scala:162 |
vecxt.matrix$Matrix.cols |
@Thin |
8 | 23% of 35 | no | matrix.scala:132 |
vecxt.matrix$Matrix.colStride |
@Thin |
8 | 23% of 35 | no | matrix.scala:138 |
vecxt.intarrays$.countsToIdx |
@HotPath |
70 | 22% of 325 | yes | intarrays.scala:244 |
vecxt.intarrays$.$minus$eq |
@AllocFree @HotPath |
67 | 21% of 325 | yes | intarrays.scala:408 |
vecxt.floatarrays$.cumsum$bang |
@AllocFree @HotPath |
27 | 8% of 325 | yes | floatarrays.scala:685 |
vecxt.doublearrays$.cumsum$bang |
@AllocFree @HotPath |
27 | 8% of 325 | yes | doublearrays.scala:950 |
vecxt.floatarrays$.$plus$eq |
@AllocFree @HotPath |
24 | 7% of 325 | no | floatarrays.scala:749 |
Method sizes
| band | methods |
|---|---|
| <= 6 (trivial, always inlined) | 1411 |
| 7-35 (inlinable cold) | 3642 |
| 36-325 (inlinable when hot) | 702 |
| 326-8000 (not inlined) | 89 |
| > 8000 (NEVER JIT COMPILED) | 0 |
| bytes | method | module | at |
|---|---|---|---|
| 5643 | CheatsheetTest$.matrixRangeSlicing |
experiments | cheatsheet.scala:118 |
| 3815 | CheatsheetTest$.matrixReverseSlicing |
experiments | cheatsheet.scala:125 |
| 3801 | CheatsheetTest$.ndArrayInt |
experiments | cheatsheet.scala:416 |
| 3779 | CheatsheetTest$.ndArrayBoolean |
experiments | cheatsheet.scala:430 |
| 3061 | CheatsheetTest$.ndArrayFloat |
experiments | cheatsheet.scala:395 |
| 3015 | CheatsheetTest$.ndArrayFloatReductions |
experiments | cheatsheet.scala:405 |
| 2811 | CheatsheetTest$.matrixCreation |
experiments | cheatsheet.scala:72 |
| 2580 | CheatsheetTest$.matrixOps |
experiments | cheatsheet.scala:167 |
| 2500 | vecxt_re.Tower.show |
vecxt_re | Tower.scala:63 |
| 1594 | vecxt.ndarrayOps$.apply |
vecxt | ndarrayOps.scala:452 |
| 1460 | CheatsheetTest$.indexingAndSlicing |
experiments | cheatsheet.scala:98 |
| 1435 | vecxt.JvmDoubleMatrix$.$plus$eq |
vecxt | doublematrix.scala:315 |
| 1420 | vecxt.JvmFloatMatrix$.floatmatrixAddScalarInPlace |
vecxt | floatmatrix.scala:442 |
| 1420 | vecxt.JvmFloatMatrix$.floatmatrixSubScalarInPlace |
vecxt | floatmatrix.scala:514 |
| 1308 | CheatsheetTest$.arrayManipulation |
experiments | cheatsheet.scala:305 |
| 1122 | vecxt.Svd$.pinv |
vecxt | svd.scala:42 |
| 1028 | CheatsheetTest$.matrixFloat |
experiments | cheatsheet.scala:445 |
| 984 | vecxt.JvmDoubleMatrix$.$plus$eq |
vecxt | doublematrix.scala:227 |
| 977 | vecxt.JvmFloatMatrix$.floatmatrixSubVector |
vecxt | floatmatrix.scala:336 |
| 969 | vecxt_re.Scenarr$.combine |
vecxt_re | scenarr.scala:160 |
| 960 | vecxt.JvmFloatMatrix$.floatmatrixAddVectorInPlace |
vecxt | floatmatrix.scala:248 |
| 960 | vecxt.JvmFloatMatrix$.floatmatrixSubVectorInPlace |
vecxt | floatmatrix.scala:360 |
| 923 | CheatsheetTest$.matrixInt |
experiments | cheatsheet.scala:461 |
| 908 | vecxt.Svd$.svd |
vecxt | svd.scala:135 |
| 896 | vecxt_re.NegativeBinomial$.volweightedMle |
vecxt_re | NegativeBinomial.scala:281 |
Proposed baseline
{
"jdkMajor": 25,
"c9": { "totalBytes": 41872, "distinctOps": 206 },
"annotated": {
"vecxt.NDArrayDoubleOps$.binaryOpGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
"vecxt.NDArrayDoubleOps$.binaryOpInPlaceGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)V": 153,
"vecxt.NDArrayDoubleOps$.compareGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
"vecxt.NDArrayDoubleOps$.compareScalarGeneral(Lvecxt/ndarray$NDArray;DLscala/Function2;)Lvecxt/ndarray$NDArray;": 157,
"vecxt.NDArrayDoubleOps$.unaryOpGeneral(Lvecxt/ndarray$NDArray;Lscala/Function1;)Lvecxt/ndarray$NDArray;": 152,
"vecxt.NDArrayFloatOps$.binaryOpGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 195,
"vecxt.NDArrayFloatOps$.binaryOpInPlaceGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)V": 162,
"vecxt.NDArrayFloatOps$.compareGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 195,
"vecxt.NDArrayFloatOps$.compareScalarGeneral(Lvecxt/ndarray$NDArray;FLscala/Function2;)Lvecxt/ndarray$NDArray;": 165,
"vecxt.NDArrayIntOps$.binaryOpGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
"vecxt.NDArrayIntOps$.binaryOpInPlaceGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)V": 153,
"vecxt.NDArrayIntOps$.compareGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
"vecxt.NDArrayIntOps$.compareScalarGeneral(Lvecxt/ndarray$NDArray;ILscala/Function2;)Lvecxt/ndarray$NDArray;": 156,
"vecxt.NDArrayIntOps$.unaryOpGeneral(Lvecxt/ndarray$NDArray;Lscala/Function1;)Lvecxt/ndarray$NDArray;": 152,
"vecxt.doublearrays$.$minus$bang([D)V": 98,
"vecxt.doublearrays$.$minus$eq([DD)V": 92,
"vecxt.doublearrays$.$minus$eq([D[D)V": 26,
"vecxt.doublearrays$.$minus([D[D)[D": 24,
"vecxt.doublearrays$.$plus$eq([DD)V": 90,
"vecxt.doublearrays$.$times$eq([D[D)V": 88,
"vecxt.doublearrays$.$times$times$bang([DD)V": 95,
"vecxt.doublearrays$.abs$bang([D)V": 98,
"vecxt.doublearrays$.abs([D)[D": 108,
"vecxt.doublearrays$.acos$bang([D)V": 98,
"vecxt.doublearrays$.acos([D)[D": 108,
"vecxt.doublearrays$.asin$bang([D)V": 98,
"vecxt.doublearrays$.asin([D)[D": 108,
"vecxt.doublearrays$.atan$bang([D)V": 98,
"vecxt.doublearrays$.atan([D)[D": 108,
"vecxt.doublearrays$.cbrt$bang([D)V": 98,
"vecxt.doublearrays$.cbrt([D)[D": 108,
"vecxt.doublearrays$.clamp$bang([DDD)V": 180,
"vecxt.doublearrays$.cos$bang([D)V": 98,
"vecxt.doublearrays$.cos([D)[D": 108,
"vecxt.doublearrays$.cosh$bang([D)V": 98,
"vecxt.doublearrays$.cosh([D)[D": 108,
"vecxt.doublearrays$.cumsum$bang([D)V": 27,
"vecxt.doublearrays$.cumsum([D)[D": 15,
"vecxt.doublearrays$.dot([D[D)D": 23,
"vecxt.doublearrays$.exp$bang([D)V": 98,
"vecxt.doublearrays$.exp([D)[D": 108,
"vecxt.doublearrays$.expm1$bang([D)V": 98,
"vecxt.doublearrays$.expm1([D)[D": 108,
"vecxt.doublearrays$.fillLinspace([DDD)V": 133,
"vecxt.doublearrays$.fma$bang([DDD)V": 85,
"vecxt.doublearrays$.increments([D)[D": 110,
"vecxt.doublearrays$.log$bang([D)V": 98,
"vecxt.doublearrays$.log([D)[D": 108,
"vecxt.doublearrays$.log10$bang([D)V": 98,
"vecxt.doublearrays$.log10([D)[D": 108,
"vecxt.doublearrays$.log1p$bang([D)V": 98,
"vecxt.doublearrays$.log1p([D)[D": 108,
"vecxt.doublearrays$.meanAndVariance([D)Lvecxt/MeanAndVariance;": 9,
"vecxt.doublearrays$.norm([D)D": 13,
"vecxt.doublearrays$.productSIMD([D)D": 86,
"vecxt.doublearrays$.sin$bang([D)V": 98,
"vecxt.doublearrays$.sin([D)[D": 108,
"vecxt.doublearrays$.sinh$bang([D)V": 98,
"vecxt.doublearrays$.sinh([D)[D": 108,
"vecxt.doublearrays$.sqrt$bang([D)V": 98,
"vecxt.doublearrays$.sqrt([D)[D": 108,
"vecxt.doublearrays$.sumSIMD([D)D": 85,
"vecxt.doublearrays$.tan$bang([D)V": 98,
"vecxt.doublearrays$.tan([D)[D": 108,
"vecxt.doublearrays$.tanh$bang([D)V": 98,
"vecxt.doublearrays$.tanh([D)[D": 108,
"vecxt.doublearrays$.unary_$minus([D)[D": 108,
"vecxt.doublearrays$.variance([DLvecxt/VarianceMode;)D": 10,
"vecxt.doublearrays.meanAndVariance([DLvecxt/VarianceMode;)Lvecxt/MeanAndVariance;": 9,
"vecxt.floatarrays$.$minus$bang([F)V": 93,
"vecxt.floatarrays$.$minus$eq([FF)V": 84,
"vecxt.floatarrays$.$minus$eq([F[F)V": 26,
"vecxt.floatarrays$.$minus([F[F)[F": 24,
"vecxt.floatarrays$.$plus$eq([FF)V": 84,
"vecxt.floatarrays$.$plus$eq([F[F)V": 24,
"vecxt.floatarrays$.$times$eq([FF)V": 78,
"vecxt.floatarrays$.$times$eq([F[F)V": 84,
"vecxt.floatarrays$.$times$times$bang([FF)V": 85,
"vecxt.floatarrays$.abs$bang([F)V": 93,
"vecxt.floatarrays$.abs([F)[F": 104,
"vecxt.floatarrays$.acos$bang([F)V": 93,
"vecxt.floatarrays$.acos([F)[F": 104,
"vecxt.floatarrays$.asin$bang([F)V": 93,
"vecxt.floatarrays$.asin([F)[F": 104,
"vecxt.floatarrays$.atan$bang([F)V": 93,
"vecxt.floatarrays$.atan([F)[F": 104,
"vecxt.floatarrays$.cbrt$bang([F)V": 93,
"vecxt.floatarrays$.cbrt([F)[F": 104,
"vecxt.floatarrays$.clamp$bang([FFF)V": 172,
"vecxt.floatarrays$.cos$bang([F)V": 93,
"vecxt.floatarrays$.cos([F)[F": 104,
"vecxt.floatarrays$.cosh$bang([F)V": 93,
"vecxt.floatarrays$.cosh([F)[F": 104,
"vecxt.floatarrays$.cumsum$bang([F)V": 27,
"vecxt.floatarrays$.cumsum([F)[F": 15,
"vecxt.floatarrays$.dot([F[F)F": 23,
"vecxt.floatarrays$.exp$bang([F)V": 93,
"vecxt.floatarrays$.exp([F)[F": 104,
"vecxt.floatarrays$.expm1$bang([F)V": 93,
"vecxt.floatarrays$.expm1([F)[F": 104,
"vecxt.floatarrays$.fma$bang([FFF)V": 80,
"vecxt.floatarrays$.increments([F)[F": 97,
"vecxt.floatarrays$.log$bang([F)V": 93,
"vecxt.floatarrays$.log([F)[F": 104,
"vecxt.floatarrays$.log10$bang([F)V": 93,
"vecxt.floatarrays$.log10([F)[F": 104,
"vecxt.floatarrays$.log1p$bang([F)V": 93,
"vecxt.floatarrays$.log1p([F)[F": 104,
"vecxt.floatarrays$.norm([F)F": 13,
"vecxt.floatarrays$.productSIMD([F)F": 82,
"vecxt.floatarrays$.sin$bang([F)V": 93,
"vecxt.floatarrays$.sin([F)[F": 104,
"vecxt.floatarrays$.sinh$bang([F)V": 93,
"vecxt.floatarrays$.sinh([F)[F": 104,
"vecxt.floatarrays$.sqrt$bang([F)V": 93,
"vecxt.floatarrays$.sqrt([F)[F": 104,
"vecxt.floatarrays$.sumSIMD([F)F": 81,
"vecxt.floatarrays$.tan$bang([F)V": 93,
"vecxt.floatarrays$.tan([F)[F": 104,
"vecxt.floatarrays$.tanh$bang([F)V": 93,
"vecxt.floatarrays$.tanh([F)[F": 104,
"vecxt.floatarrays$.unary_$minus([F)[F": 104,
"vecxt.intarrays$.$bang$colon$eq([II)[Z": 94,
"vecxt.intarrays$.$bang$colon$eq([I[I)[Z": 106,
"vecxt.intarrays$.$eq$colon$eq([II)[Z": 94,
"vecxt.intarrays$.$eq$colon$eq([I[I)[Z": 106,
"vecxt.intarrays$.$greater$eq([II)[Z": 94,
"vecxt.intarrays$.$greater$eq([I[I)[Z": 106,
"vecxt.intarrays$.$greater([II)[Z": 94,
"vecxt.intarrays$.$greater([I[I)[Z": 106,
"vecxt.intarrays$.$less$eq([II)[Z": 94,
"vecxt.intarrays$.$less$eq([I[I)[Z": 106,
"vecxt.intarrays$.$less([II)[Z": 94,
"vecxt.intarrays$.$less([I[I)[Z": 106,
"vecxt.intarrays$.$minus$eq([II)V": 67,
"vecxt.intarrays$.$minus$eq([I[I)V": 84,
"vecxt.intarrays$.$minus([II)[I": 16,
"vecxt.intarrays$.$minus([I[I)[I": 24,
"vecxt.intarrays$.$plus$eq([I[I)V": 84,
"vecxt.intarrays$.$plus([I[I)[I": 24,
"vecxt.intarrays$.countsToIdx([I)[I": 70,
"vecxt.intarrays$.dot([I[I)I": 108,
"vecxt.intarrays$.increments([I)[I": 121,
"vecxt.intarrays$.maxSIMD([I)I": 82,
"vecxt.intarrays$.mean([I)D": 11,
"vecxt.intarrays$.meanAndVariance([I)Lvecxt/MeanAndVariance;": 9,
"vecxt.intarrays$.minSIMD([I)I": 82,
"vecxt.intarrays$.std([I)D": 9,
"vecxt.intarrays$.std([ILvecxt/VarianceMode;)D": 10,
"vecxt.intarrays$.stdDev([I)D": 9,
"vecxt.intarrays$.sumSIMD([I)I": 81,
"vecxt.intarrays$.variance([I)D": 9,
"vecxt.intarrays$.variance([ILvecxt/VarianceMode;)D": 10,
"vecxt.intarrays.gt([II)[Z": 9,
"vecxt.intarrays.gt([I[I)[Z": 9,
"vecxt.intarrays.gte([II)[Z": 9,
"vecxt.intarrays.gte([I[I)[Z": 9,
"vecxt.intarrays.lt([II)[Z": 9,
"vecxt.intarrays.lt([I[I)[Z": 9,
"vecxt.intarrays.lte([II)[Z": 9,
"vecxt.intarrays.lte([I[I)[Z": 9,
"vecxt.intarrays.meanAndVariance([ILvecxt/VarianceMode;)Lvecxt/MeanAndVariance;": 9,
"vecxt.intarrays.stdDev([ILvecxt/VarianceMode;)D": 9,
"vecxt.matrix$Layout.linearIndex(II)I": 19,
"vecxt.matrix$Matrix.colStride()I": 8,
"vecxt.matrix$Matrix.cols()I": 8,
"vecxt.matrix$Matrix.hasSimpleContiguousMemoryLayout()Z": 8,
"vecxt.matrix$Matrix.isDenseColMajor()Z": 8,
"vecxt.matrix$Matrix.isDenseRowMajor()Z": 8,
"vecxt.matrix$Matrix.numel()I": 8,
"vecxt.matrix$Matrix.offset()I": 8,
"vecxt.matrix$Matrix.rowStride()I": 8,
"vecxt.matrix$Matrix.rows()I": 8,
"vecxt.ndarray$.mkNDArray(Ljava/lang/Object;[I[II)Lvecxt/ndarray$NDArray;": 13,
"vecxt.ndarray.shapeArray(Lvecxt/ndarray$NDArray;)[I": 8,
"vecxt.ndarrayOps.expandDims(Lvecxt/ndarray$NDArray;I)Lvecxt/ndarray$NDArray;": 9
}
}
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Six methods per file, in doublearrays and its floatarrays mirror. None of them meets any of the three conditions the inlining policy keeps
inlinefor: no closure parameter, no VectorOperators constant to hold, and a concrete element type at every point they touch an element.cumsum! @hotpath @AllocFree
cumsum @thin
dot @thin
norm @thin
-= @thin (the Array overload; see below)
cumsum!is the interesting one. A prefix sum carries a dependency —vec(i)needs the value just written tovec(i - 1)— so it cannot be vectorised and there is no Vector API in the body at all. That makes it the only@AllocFreekernel in the library whose zero is unconditional rather than contingent on C2 applying intrinsics: there is no DoubleVector whose failure to scalarise could put bytes on the heap. It is worth measuring precisely because it is the control case for every other kernel in D1 — if it ever fails, the harness is wrong and not the kernel. Added to D1Suite and D1EAOffSuite for both element types, which brings both suites to 31 tests and keeps every @AllocFree method covered in both.dot,normand the array-argument-=are@Thinrather than@HotPath, and the distinction is the annotation's own wording:@HotPathdescribes "code that runs once per element", and the per-element loop in these is inside netlib's ddot/dnrm2/daxpy, not in the method. What the method does is name the operation and dispatch, which is@Thin's definition. C3's no-backward-branch assertion holds for the same reason and would fail if anyone open-coded the loop back in, which is the right outcome.None of the three is
@AllocFree.blasisJavaBLAS.getInstance, so these call into a third-party pure-Java implementation whose allocation behaviour is not visible from here and has never been measured. Both annotations that turned out to be false —**!andintarrays.dot— were applied by inspection, so inspection is not the standard being used here.Worth a comment where it lands:
-=(Array)is a BLAS forwarder while-=(Double)is a hand-written Vector API loop carrying @hotpath @AllocFree. Same name, different implementations, so different annotations, and not an inconsistency for someone to tidy up later.Scope note: four of these were named directly;
normand-were added because they sit inside the same block and are the same two shapes, and leaving theminlinewould half-convert it.add,+and+=in the same file are the same shapes again but are left alone —addforwards to+, so it cannot become a clean forwarder until+is emitted, and that ordering is a batch of its own.