Skip to content

Doc machine code budgets - #121

Merged
Quafadas merged 5 commits into
mainfrom
doc-machine-code-budgets
Aug 4, 2026
Merged

Doc machine code budgets#121
Quafadas merged 5 commits into
mainfrom
doc-machine-code-budgets

Conversation

@Quafadas

@Quafadas Quafadas commented Aug 4, 2026

Copy link
Copy Markdown
Owner

No description provided.

Simon Parten and others added 3 commits August 4, 2026 11:56
…ode budget

Doc-only. A LogCompilation probe of the jitAudit kernels (Microsoft OpenJDK
25.0.4, x86-64, 8-lane double) contradicts three claims the repo currently makes.

The measurements, reading `stub_offset - insts_offset` off each c2 <nmethod>:

  Object.<init>                       1 bytecode  ->  208 bytes machine code
  doublearrays.clamp!               180           -> 1728   (9.6x,   69% of 2500)
  doublearrays.meanAndVarianceTwoPass 248         -> 1688   (6.8x,   68%)
  vecxt.all.clamp!  (export forwarder) 11         -> 1696   (154x,   68%)

~200 bytes of fixed overhead plus 7-10x the bytecode for vectorised code with a
masked tail. At that ratio `InlineSmallCode` (2500, machine code) is reached at
roughly 260-300 bytecodes — *below* `FreqInlineSize` (325, bytecode). For a SIMD
kernel the machine-code limit binds first.

HotPath.java claimed check C2 meant "C2 will inline it into its callers once it is
hot". It does not: that is a necessary condition, not a sufficient one, and the
doc now says so and points at D2 for the compiled size.

Thin.java asserted a 35-bytecode budget as the thing that makes the public API
zero-cost at a cold call site. `vecxt.all.clamp!` satisfies it by a factor of
three while being 1696 bytes of machine code. Forwarders are where the ratio is
most extreme, precisely because the body they forward to gets pulled in — and the
`vecxt.all` forwarders are excluded from the baseline by `primaryAnnotated`, so
their compiled size is unmeasured twice over.

The blog gets a new subsection under the threshold table, "The limit that is not
measured in bytecodes", because every threshold in that table is a bytecode count
and this one is not. Also records that `FreqInlineSize` and `InlineSmallCode` are
`C2 pd product` — platform-dependent — so two of the four budgets the page relies
on are properties of the machine rather than of HotSpot. And that C1/C2/C3 read
bytecode and therefore cannot see any of this.

`doublearrays.variance(mode)` gets the specific consequence written down: its
`@AllocFree` zero depends on C2 inlining `meanAndVarianceTwoPass`, which is at 68%
of the limit. If that crosses, the pair escapes and the symptom is a D1 failure
complaining about allocation rather than about inlining.

Two things deliberately not claimed. The ratio is one workload on one CPU at one
lane width, so the direction is established and the crossover point approximate.
And whether HotSpot's MaxTrivialSize/MaxInlineSize fast paths let a small callee
bypass the InlineSmallCode veto is unverified — both docs say so rather than
guessing, because it decides whether the forwarder finding is a curiosity or a
hazard.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tch @AllocFree

Three things, one story: the LogCompilation-parsing checks are not being built,
so the annotations and docs that pointed at them as future work are corrected, and
the check that does cover the same ground gets its population widened.

== D2 and D5 will not be built ==

The probe settled feasibility and then undermined the case. <intrinsic> is
emitted and self-describing, and a Mill test fork produces a complete log — so D2
is buildable. But the vocabulary is JDK-internal with no compatibility contract
(the plan's guessed inline_fail strings were wrong in five places), and the drift
is indistinguishable from the regression: a renamed intrinsic id shrinks the
observed set, which is exactly what losing vectorisation looks like, with no
invariant to violate and no cross-check available.

The line that fell out of it is which API a check depends on. D1 uses
ThreadMXBean, D3 uses VectorSpecies, D4 uses arithmetic, eaOffTest uses a
product-grade -XX: flag — all public and contractual. D2 and D5 would have
depended on the compiler's diagnostic output instead. That is the dividing line,
not the amount of work.

Recorded in jitAudit/package.mill rather than left as an absence, because an
unexplained gap in a checklist reads as an oversight. The forward references
written in the previous commit — HotPath.java's "check D2's business", the
variance comment's "see check D2", the blog's pointer — now say the limit is
documented and unenforced, and name what covers it indirectly.

== eaOffTest is what replaces D2, and it is switched off ==

Reading it properly while answering "do these modules earn their keep": D1 with
EA enabled cannot distinguish "the SIMD intrinsics were applied" from "they were
not, and EA scalarised the software-path objects instead" — both read as zero.
The EA-off scope can, because on the software path with EA off the objects reach
the heap. That is D2's question answered from an observable consequence rather
than from XML, it is already written, and its CI step is commented out. Its
scaladoc now says so. Nothing removed.

== D1 widened from 20 to 29 kernels ==

Every @AllocFree method now has a test behind it. The nine added:

  floatarrays  clamp!, +=(Array[Float]), *=(Array[Float]), *=(Float)
  intarrays    minSIMD, maxSIMD, +=(Array[Int]), -=(Array[Int]), -=(Int)

An annotation with no test is an assertion nobody has checked, which is how
intarrays.dot carried @AllocFree while allocating a dead array per call. The
class doc now states that coverage is every annotated method and that the two are
kept in step by hand — and drops the brittle counts that have needed correcting
twice.

Also notes why the kernel lists cannot be factored into a shared collection:
assertAllocFree must stay inline so each call site gets a monomorphic measurement
loop, and driving them from a List[() => Unit] would reintroduce the megamorphic
Function0.apply() dispatch that stops C2 inlining through to the Vector API calls.
The duplication is load-bearing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown

Bytecode audit (Tier 1) — ⚠️ no failures, 1 WARN

JDK 25.0.1 (major 25), OpenJDK 64-Bit Server VM

Vector lanes (DoubleVector.SPECIES_PREFERRED.length()): 8

threshold value provenance
MaxTrivialSize 6 discovered
MaxInlineSize 35 discovered
FreqInlineSize 325 discovered
MaxInlineLevel 15 discovered
InlineSmallCode 2500 discovered
NodeCountInliningCutoff 18000 assumed
HugeMethodLimit 8000 assumed

-XX:HugeMethodLimit= was rejected on the command line: a develop flag compiled out of this product build, so 8000 is taken from the HotSpot source and cannot be confirmed against the running JVM.

metric now baseline delta
cheatsheet methods 97
total bytes 42067 51103 -17.7%
distinct library ops 198 178 +11.2%
bytes per op 212.5 287.1 -26.0%
severity check at method detail
WARN C1 cheatsheet.scala:118 CheatsheetTest$.matrixRangeSlicing 5643 bytes, past 68% of HugeMethodLimit=8000 (assumed)
Annotated methods (164)
method annotations bytes budget used loop at
vecxt.intarrays$.$plus @Thin 24 69% of 35 no intarrays.scala:546
vecxt.intarrays$.$minus @Thin 24 69% of 35 no intarrays.scala:399
vecxt.NDArrayFloatOps$.compareGeneral @HotPath 195 60% of 325 yes ndarrayFloatOps.scala:67
vecxt.NDArrayFloatOps$.binaryOpGeneral @HotPath 195 60% of 325 yes ndarrayFloatOps.scala:20
vecxt.NDArrayIntOps$.compareGeneral @HotPath 186 57% of 325 yes ndarrayIntOps.scala:67
vecxt.NDArrayIntOps$.binaryOpGeneral @HotPath 186 57% of 325 yes ndarrayIntOps.scala:19
vecxt.NDArrayDoubleOps$.compareGeneral @HotPath 186 57% of 325 yes ndarrayDoubleOps.scala:74
vecxt.NDArrayDoubleOps$.binaryOpGeneral @HotPath 186 57% of 325 yes ndarrayDoubleOps.scala:24
vecxt.doublearrays$.clamp$bang @AllocFree @HotPath 180 55% of 325 yes doublearrays.scala:872
vecxt.matrix$Layout.linearIndex @Thin 19 54% of 35 no matrix.scala:48
vecxt.floatarrays$.clamp$bang @AllocFree @HotPath 172 53% of 325 yes floatarrays.scala:417
vecxt.NDArrayFloatOps$.compareScalarGeneral @HotPath 165 51% of 325 yes ndarrayFloatOps.scala:94
vecxt.NDArrayFloatOps$.binaryOpInPlaceGeneral @HotPath 162 50% of 325 yes ndarrayFloatOps.scala:119
vecxt.NDArrayDoubleOps$.compareScalarGeneral @HotPath 157 48% of 325 yes ndarrayDoubleOps.scala:102
vecxt.NDArrayIntOps$.compareScalarGeneral @HotPath 156 48% of 325 yes ndarrayIntOps.scala:94
vecxt.NDArrayIntOps$.binaryOpInPlaceGeneral @HotPath 153 47% of 325 yes ndarrayIntOps.scala:119
vecxt.NDArrayDoubleOps$.binaryOpInPlaceGeneral @HotPath 153 47% of 325 yes ndarrayDoubleOps.scala:131
vecxt.NDArrayIntOps$.unaryOpGeneral @HotPath 152 47% of 325 yes ndarrayIntOps.scala:42
vecxt.NDArrayDoubleOps$.unaryOpGeneral @HotPath 152 47% of 325 yes ndarrayDoubleOps.scala:48
vecxt.intarrays$.$minus @Thin 16 46% of 35 no intarrays.scala:519
vecxt.doublearrays$.fillLinspace @AllocFree @HotPath 133 41% of 325 yes doublearrays.scala:34
vecxt.intarrays$.increments @HotPath 121 37% of 325 yes intarrays.scala:219
vecxt.ndarray$.mkNDArray @Thin 13 37% of 35 no ndarray.scala:188
vecxt.doublearrays$.increments @HotPath 110 34% of 325 yes doublearrays.scala:397
vecxt.intarrays$.dot @AllocFree @HotPath 108 33% of 325 yes intarrays.scala:372
vecxt.doublearrays$.unary_$minus @HotPath 108 33% of 325 yes doublearrays.scala:172
vecxt.doublearrays$.tanh @HotPath 108 33% of 325 yes doublearrays.scala:365
vecxt.doublearrays$.tan @HotPath 108 33% of 325 yes doublearrays.scala:354
vecxt.doublearrays$.sqrt @HotPath 108 33% of 325 yes doublearrays.scala:317
vecxt.doublearrays$.sinh @HotPath 108 33% of 325 yes doublearrays.scala:339
vecxt.doublearrays$.sin @HotPath 108 33% of 325 yes doublearrays.scala:328
vecxt.doublearrays$.log1p @HotPath 108 33% of 325 yes doublearrays.scala:306
vecxt.doublearrays$.log10 @HotPath 108 33% of 325 yes doublearrays.scala:295
vecxt.doublearrays$.log @HotPath 108 33% of 325 yes doublearrays.scala:284
vecxt.doublearrays$.expm1 @HotPath 108 33% of 325 yes doublearrays.scala:273
vecxt.doublearrays$.exp @HotPath 108 33% of 325 yes doublearrays.scala:262
vecxt.doublearrays$.cosh @HotPath 108 33% of 325 yes doublearrays.scala:251
vecxt.doublearrays$.cos @HotPath 108 33% of 325 yes doublearrays.scala:240
vecxt.doublearrays$.cbrt @HotPath 108 33% of 325 yes doublearrays.scala:229
vecxt.doublearrays$.atan @HotPath 108 33% of 325 yes doublearrays.scala:218
vecxt.doublearrays$.asin @HotPath 108 33% of 325 yes doublearrays.scala:207
vecxt.doublearrays$.acos @HotPath 108 33% of 325 yes doublearrays.scala:196
vecxt.doublearrays$.abs @HotPath 108 33% of 325 yes doublearrays.scala:184
vecxt.intarrays$.$less @HotPath 106 33% of 325 yes intarrays.scala:48
vecxt.intarrays$.$less$eq @HotPath 106 33% of 325 yes intarrays.scala:52
vecxt.intarrays$.$greater @HotPath 106 33% of 325 yes intarrays.scala:56
vecxt.intarrays$.$greater$eq @HotPath 106 33% of 325 yes intarrays.scala:60
vecxt.intarrays$.$eq$colon$eq @HotPath 106 33% of 325 yes intarrays.scala:40
vecxt.intarrays$.$bang$colon$eq @HotPath 106 33% of 325 yes intarrays.scala:44
vecxt.floatarrays$.unary_$minus @HotPath 104 32% of 325 yes floatarrays.scala:114
vecxt.floatarrays$.tanh @HotPath 104 32% of 325 yes floatarrays.scala:303
vecxt.floatarrays$.tan @HotPath 104 32% of 325 yes floatarrays.scala:292
vecxt.floatarrays$.sqrt @HotPath 104 32% of 325 yes floatarrays.scala:259
vecxt.floatarrays$.sinh @HotPath 104 32% of 325 yes floatarrays.scala:281
vecxt.floatarrays$.sin @HotPath 104 32% of 325 yes floatarrays.scala:270
vecxt.floatarrays$.log1p @HotPath 104 32% of 325 yes floatarrays.scala:248
vecxt.floatarrays$.log10 @HotPath 104 32% of 325 yes floatarrays.scala:237
vecxt.floatarrays$.log @HotPath 104 32% of 325 yes floatarrays.scala:226
vecxt.floatarrays$.expm1 @HotPath 104 32% of 325 yes floatarrays.scala:215
vecxt.floatarrays$.exp @HotPath 104 32% of 325 yes floatarrays.scala:204
vecxt.floatarrays$.cosh @HotPath 104 32% of 325 yes floatarrays.scala:193
vecxt.floatarrays$.cos @HotPath 104 32% of 325 yes floatarrays.scala:182
vecxt.floatarrays$.cbrt @HotPath 104 32% of 325 yes floatarrays.scala:171
vecxt.floatarrays$.atan @HotPath 104 32% of 325 yes floatarrays.scala:160
vecxt.floatarrays$.asin @HotPath 104 32% of 325 yes floatarrays.scala:149
vecxt.floatarrays$.acos @HotPath 104 32% of 325 yes floatarrays.scala:138
vecxt.floatarrays$.abs @HotPath 104 32% of 325 yes floatarrays.scala:126
vecxt.intarrays$.mean @Thin 11 31% of 35 no intarrays.scala:287
vecxt.doublearrays$.tanh$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.tan$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.sqrt$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.sinh$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.sin$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.log1p$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.log10$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.log$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.expm1$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.exp$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.cosh$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.cos$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.cbrt$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.atan$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.asin$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.acos$bang @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.abs$bang @AllocFree @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.doublearrays$.$minus$bang @AllocFree @HotPath 98 30% of 325 yes doublearrays.scala:151
vecxt.floatarrays$.increments @HotPath 97 30% of 325 yes floatarrays.scala:652
vecxt.doublearrays$.$times$times$bang @HotPath 95 29% of 325 yes doublearrays.scala:372
vecxt.intarrays$.$less @HotPath 94 29% of 325 yes intarrays.scala:139
vecxt.intarrays$.$less$eq @HotPath 94 29% of 325 yes intarrays.scala:143
vecxt.intarrays$.$greater @HotPath 94 29% of 325 yes intarrays.scala:147
vecxt.intarrays$.$greater$eq @HotPath 94 29% of 325 yes intarrays.scala:151
vecxt.intarrays$.$eq$colon$eq @HotPath 94 29% of 325 yes intarrays.scala:131
vecxt.intarrays$.$bang$colon$eq @HotPath 94 29% of 325 yes intarrays.scala:135
vecxt.floatarrays$.tanh$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.tan$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.sqrt$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.sinh$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.sin$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.log1p$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.log10$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.log$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.expm1$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.exp$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.cosh$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.cos$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.cbrt$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.atan$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.asin$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.acos$bang @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.abs$bang @AllocFree @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.floatarrays$.$minus$bang @AllocFree @HotPath 93 29% of 325 yes floatarrays.scala:94
vecxt.intarrays$.variance @Thin 10 29% of 35 no intarrays.scala:304
vecxt.intarrays$.std @Thin 10 29% of 35 no intarrays.scala:361
vecxt.doublearrays$.variance @AllocFree @Thin 10 29% of 35 no doublearrays.scala:541
vecxt.doublearrays$.$minus$eq @AllocFree @HotPath 92 28% of 325 yes doublearrays.scala:1063
vecxt.doublearrays$.$plus$eq @AllocFree @HotPath 90 28% of 325 yes doublearrays.scala:998
vecxt.doublearrays$.$times$eq @AllocFree @HotPath 88 27% of 325 yes doublearrays.scala:1135
vecxt.doublearrays$.productSIMD @AllocFree @HotPath 86 26% of 325 yes doublearrays.scala:700
vecxt.floatarrays$.$times$times$bang @HotPath 85 26% of 325 yes floatarrays.scala:314
vecxt.doublearrays$.sumSIMD @AllocFree @HotPath 85 26% of 325 yes doublearrays.scala:677
vecxt.doublearrays$.fma$bang @AllocFree @HotPath 85 26% of 325 yes doublearrays.scala:1038
vecxt.intarrays$.$plus$eq @AllocFree @HotPath 84 26% of 325 yes intarrays.scala:555
vecxt.intarrays$.$minus$eq @AllocFree @HotPath 84 26% of 325 yes intarrays.scala:527
vecxt.floatarrays$.$times$eq @AllocFree @HotPath 84 26% of 325 yes floatarrays.scala:874
vecxt.floatarrays$.$plus$eq @AllocFree @HotPath 84 26% of 325 yes floatarrays.scala:772
vecxt.floatarrays$.$minus$eq @AllocFree @HotPath 84 26% of 325 yes floatarrays.scala:812
vecxt.ndarrayOps.expandDims @Thin 9 26% of 35 no ndarrayOps.scala
vecxt.intarrays.stdDev @Thin 9 26% of 35 no intarrays.scala
vecxt.intarrays.meanAndVariance @Thin 9 26% of 35 no intarrays.scala
vecxt.intarrays.lte @Thin 9 26% of 35 no intarrays.scala
vecxt.intarrays.lte @Thin 9 26% of 35 no intarrays.scala
vecxt.intarrays.lt @Thin 9 26% of 35 no intarrays.scala
vecxt.intarrays.lt @Thin 9 26% of 35 no intarrays.scala
vecxt.intarrays.gte @Thin 9 26% of 35 no intarrays.scala
vecxt.intarrays.gte @Thin 9 26% of 35 no intarrays.scala
vecxt.intarrays.gt @Thin 9 26% of 35 no intarrays.scala
vecxt.intarrays.gt @Thin 9 26% of 35 no intarrays.scala
vecxt.intarrays$.variance @Thin 9 26% of 35 no intarrays.scala:291
vecxt.intarrays$.stdDev @Thin 9 26% of 35 no intarrays.scala:364
vecxt.intarrays$.std @Thin 9 26% of 35 no intarrays.scala:357
vecxt.intarrays$.meanAndVariance @Thin 9 26% of 35 no intarrays.scala:308
vecxt.doublearrays.meanAndVariance @Thin 9 26% of 35 no doublearrays.scala
vecxt.doublearrays$.meanAndVariance @Thin 9 26% of 35 no doublearrays.scala:555
vecxt.intarrays$.minSIMD @AllocFree @HotPath 82 25% of 325 yes intarrays.scala:575
vecxt.intarrays$.maxSIMD @AllocFree @HotPath 82 25% of 325 yes intarrays.scala:595
vecxt.floatarrays$.productSIMD @AllocFree @HotPath 82 25% of 325 yes floatarrays.scala:537
vecxt.intarrays$.sumSIMD @AllocFree @HotPath 81 25% of 325 yes intarrays.scala:265
vecxt.floatarrays$.sumSIMD @AllocFree @HotPath 81 25% of 325 yes floatarrays.scala:516
vecxt.floatarrays$.fma$bang @AllocFree @HotPath 80 25% of 325 yes floatarrays.scala:340
vecxt.floatarrays$.$times$eq @AllocFree @HotPath 78 24% of 325 yes floatarrays.scala:926
vecxt.ndarray.shapeArray @Thin 8 23% of 35 no ndarray.scala
vecxt.matrix$Matrix.rows @Thin 8 23% of 35 no matrix.scala:129
vecxt.matrix$Matrix.rowStride @Thin 8 23% of 35 no matrix.scala:135
vecxt.matrix$Matrix.offset @Thin 8 23% of 35 no matrix.scala:141
vecxt.matrix$Matrix.numel @Thin 8 23% of 35 no matrix.scala:144
vecxt.matrix$Matrix.isDenseRowMajor @Thin 8 23% of 35 no matrix.scala:150
vecxt.matrix$Matrix.isDenseColMajor @Thin 8 23% of 35 no matrix.scala:147
vecxt.matrix$Matrix.hasSimpleContiguousMemoryLayout @Thin 8 23% of 35 no matrix.scala:162
vecxt.matrix$Matrix.cols @Thin 8 23% of 35 no matrix.scala:132
vecxt.matrix$Matrix.colStride @Thin 8 23% of 35 no matrix.scala:138
vecxt.intarrays$.countsToIdx @HotPath 70 22% of 325 yes intarrays.scala:244
vecxt.intarrays$.$minus$eq @AllocFree @HotPath 67 21% of 325 yes intarrays.scala:408
vecxt.floatarrays$.$plus$eq @AllocFree @HotPath 24 7% of 325 no floatarrays.scala:731
Method sizes
band methods
<= 6 (trivial, always inlined) 1411
7-35 (inlinable cold) 3569
36-325 (inlinable when hot) 702
326-8000 (not inlined) 90
> 8000 (NEVER JIT COMPILED) 0
bytes method module at
5643 CheatsheetTest$.matrixRangeSlicing experiments cheatsheet.scala:118
3815 CheatsheetTest$.matrixReverseSlicing experiments cheatsheet.scala:125
3801 CheatsheetTest$.ndArrayInt experiments cheatsheet.scala:416
3779 CheatsheetTest$.ndArrayBoolean experiments cheatsheet.scala:430
3061 CheatsheetTest$.ndArrayFloat experiments cheatsheet.scala:395
3015 CheatsheetTest$.ndArrayFloatReductions experiments cheatsheet.scala:405
2811 CheatsheetTest$.matrixCreation experiments cheatsheet.scala:72
2597 CheatsheetTest$.matrixOps experiments cheatsheet.scala:167
2500 vecxt_re.Tower.show vecxt_re Tower.scala:63
1594 vecxt.ndarrayOps$.apply vecxt ndarrayOps.scala:452
1460 CheatsheetTest$.indexingAndSlicing experiments cheatsheet.scala:98
1435 vecxt.JvmDoubleMatrix$.$plus$eq vecxt doublematrix.scala:315
1420 vecxt.JvmFloatMatrix$.floatmatrixAddScalarInPlace vecxt floatmatrix.scala:442
1420 vecxt.JvmFloatMatrix$.floatmatrixSubScalarInPlace vecxt floatmatrix.scala:514
1308 CheatsheetTest$.arrayManipulation experiments cheatsheet.scala:305
1122 vecxt.Svd$.pinv vecxt svd.scala:42
1040 vecxt.JvmFloatMatrix$.floatmatrixSubVector vecxt floatmatrix.scala:336
1028 CheatsheetTest$.matrixFloat experiments cheatsheet.scala:445
984 vecxt.JvmDoubleMatrix$.$plus$eq vecxt doublematrix.scala:227
969 vecxt_re.Scenarr$.combine vecxt_re scenarr.scala:160
960 vecxt.JvmFloatMatrix$.floatmatrixAddVectorInPlace vecxt floatmatrix.scala:248
960 vecxt.JvmFloatMatrix$.floatmatrixSubVectorInPlace vecxt floatmatrix.scala:360
923 CheatsheetTest$.matrixInt experiments cheatsheet.scala:461
908 vecxt.Svd$.svd vecxt svd.scala:135
896 vecxt_re.NegativeBinomial$.volweightedMle vecxt_re NegativeBinomial.scala:281
Proposed baseline
{
  "jdkMajor": 25,
  "c9": { "totalBytes": 42067, "distinctOps": 198 },
  "annotated": {
    "vecxt.NDArrayDoubleOps$.binaryOpGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
    "vecxt.NDArrayDoubleOps$.binaryOpInPlaceGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)V": 153,
    "vecxt.NDArrayDoubleOps$.compareGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
    "vecxt.NDArrayDoubleOps$.compareScalarGeneral(Lvecxt/ndarray$NDArray;DLscala/Function2;)Lvecxt/ndarray$NDArray;": 157,
    "vecxt.NDArrayDoubleOps$.unaryOpGeneral(Lvecxt/ndarray$NDArray;Lscala/Function1;)Lvecxt/ndarray$NDArray;": 152,
    "vecxt.NDArrayFloatOps$.binaryOpGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 195,
    "vecxt.NDArrayFloatOps$.binaryOpInPlaceGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)V": 162,
    "vecxt.NDArrayFloatOps$.compareGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 195,
    "vecxt.NDArrayFloatOps$.compareScalarGeneral(Lvecxt/ndarray$NDArray;FLscala/Function2;)Lvecxt/ndarray$NDArray;": 165,
    "vecxt.NDArrayIntOps$.binaryOpGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
    "vecxt.NDArrayIntOps$.binaryOpInPlaceGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)V": 153,
    "vecxt.NDArrayIntOps$.compareGeneral(Lvecxt/ndarray$NDArray;Lvecxt/ndarray$NDArray;Lscala/Function2;)Lvecxt/ndarray$NDArray;": 186,
    "vecxt.NDArrayIntOps$.compareScalarGeneral(Lvecxt/ndarray$NDArray;ILscala/Function2;)Lvecxt/ndarray$NDArray;": 156,
    "vecxt.NDArrayIntOps$.unaryOpGeneral(Lvecxt/ndarray$NDArray;Lscala/Function1;)Lvecxt/ndarray$NDArray;": 152,
    "vecxt.doublearrays$.$minus$bang([D)V": 98,
    "vecxt.doublearrays$.$minus$eq([DD)V": 92,
    "vecxt.doublearrays$.$plus$eq([DD)V": 90,
    "vecxt.doublearrays$.$times$eq([D[D)V": 88,
    "vecxt.doublearrays$.$times$times$bang([DD)V": 95,
    "vecxt.doublearrays$.abs$bang([D)V": 98,
    "vecxt.doublearrays$.abs([D)[D": 108,
    "vecxt.doublearrays$.acos$bang([D)V": 98,
    "vecxt.doublearrays$.acos([D)[D": 108,
    "vecxt.doublearrays$.asin$bang([D)V": 98,
    "vecxt.doublearrays$.asin([D)[D": 108,
    "vecxt.doublearrays$.atan$bang([D)V": 98,
    "vecxt.doublearrays$.atan([D)[D": 108,
    "vecxt.doublearrays$.cbrt$bang([D)V": 98,
    "vecxt.doublearrays$.cbrt([D)[D": 108,
    "vecxt.doublearrays$.clamp$bang([DDD)V": 180,
    "vecxt.doublearrays$.cos$bang([D)V": 98,
    "vecxt.doublearrays$.cos([D)[D": 108,
    "vecxt.doublearrays$.cosh$bang([D)V": 98,
    "vecxt.doublearrays$.cosh([D)[D": 108,
    "vecxt.doublearrays$.exp$bang([D)V": 98,
    "vecxt.doublearrays$.exp([D)[D": 108,
    "vecxt.doublearrays$.expm1$bang([D)V": 98,
    "vecxt.doublearrays$.expm1([D)[D": 108,
    "vecxt.doublearrays$.fillLinspace([DDD)V": 133,
    "vecxt.doublearrays$.fma$bang([DDD)V": 85,
    "vecxt.doublearrays$.increments([D)[D": 110,
    "vecxt.doublearrays$.log$bang([D)V": 98,
    "vecxt.doublearrays$.log([D)[D": 108,
    "vecxt.doublearrays$.log10$bang([D)V": 98,
    "vecxt.doublearrays$.log10([D)[D": 108,
    "vecxt.doublearrays$.log1p$bang([D)V": 98,
    "vecxt.doublearrays$.log1p([D)[D": 108,
    "vecxt.doublearrays$.meanAndVariance([D)Lvecxt/MeanAndVariance;": 9,
    "vecxt.doublearrays$.productSIMD([D)D": 86,
    "vecxt.doublearrays$.sin$bang([D)V": 98,
    "vecxt.doublearrays$.sin([D)[D": 108,
    "vecxt.doublearrays$.sinh$bang([D)V": 98,
    "vecxt.doublearrays$.sinh([D)[D": 108,
    "vecxt.doublearrays$.sqrt$bang([D)V": 98,
    "vecxt.doublearrays$.sqrt([D)[D": 108,
    "vecxt.doublearrays$.sumSIMD([D)D": 85,
    "vecxt.doublearrays$.tan$bang([D)V": 98,
    "vecxt.doublearrays$.tan([D)[D": 108,
    "vecxt.doublearrays$.tanh$bang([D)V": 98,
    "vecxt.doublearrays$.tanh([D)[D": 108,
    "vecxt.doublearrays$.unary_$minus([D)[D": 108,
    "vecxt.doublearrays$.variance([DLvecxt/VarianceMode;)D": 10,
    "vecxt.doublearrays.meanAndVariance([DLvecxt/VarianceMode;)Lvecxt/MeanAndVariance;": 9,
    "vecxt.floatarrays$.$minus$bang([F)V": 93,
    "vecxt.floatarrays$.$minus$eq([FF)V": 84,
    "vecxt.floatarrays$.$plus$eq([FF)V": 84,
    "vecxt.floatarrays$.$plus$eq([F[F)V": 24,
    "vecxt.floatarrays$.$times$eq([FF)V": 78,
    "vecxt.floatarrays$.$times$eq([F[F)V": 84,
    "vecxt.floatarrays$.$times$times$bang([FF)V": 85,
    "vecxt.floatarrays$.abs$bang([F)V": 93,
    "vecxt.floatarrays$.abs([F)[F": 104,
    "vecxt.floatarrays$.acos$bang([F)V": 93,
    "vecxt.floatarrays$.acos([F)[F": 104,
    "vecxt.floatarrays$.asin$bang([F)V": 93,
    "vecxt.floatarrays$.asin([F)[F": 104,
    "vecxt.floatarrays$.atan$bang([F)V": 93,
    "vecxt.floatarrays$.atan([F)[F": 104,
    "vecxt.floatarrays$.cbrt$bang([F)V": 93,
    "vecxt.floatarrays$.cbrt([F)[F": 104,
    "vecxt.floatarrays$.clamp$bang([FFF)V": 172,
    "vecxt.floatarrays$.cos$bang([F)V": 93,
    "vecxt.floatarrays$.cos([F)[F": 104,
    "vecxt.floatarrays$.cosh$bang([F)V": 93,
    "vecxt.floatarrays$.cosh([F)[F": 104,
    "vecxt.floatarrays$.exp$bang([F)V": 93,
    "vecxt.floatarrays$.exp([F)[F": 104,
    "vecxt.floatarrays$.expm1$bang([F)V": 93,
    "vecxt.floatarrays$.expm1([F)[F": 104,
    "vecxt.floatarrays$.fma$bang([FFF)V": 80,
    "vecxt.floatarrays$.increments([F)[F": 97,
    "vecxt.floatarrays$.log$bang([F)V": 93,
    "vecxt.floatarrays$.log([F)[F": 104,
    "vecxt.floatarrays$.log10$bang([F)V": 93,
    "vecxt.floatarrays$.log10([F)[F": 104,
    "vecxt.floatarrays$.log1p$bang([F)V": 93,
    "vecxt.floatarrays$.log1p([F)[F": 104,
    "vecxt.floatarrays$.productSIMD([F)F": 82,
    "vecxt.floatarrays$.sin$bang([F)V": 93,
    "vecxt.floatarrays$.sin([F)[F": 104,
    "vecxt.floatarrays$.sinh$bang([F)V": 93,
    "vecxt.floatarrays$.sinh([F)[F": 104,
    "vecxt.floatarrays$.sqrt$bang([F)V": 93,
    "vecxt.floatarrays$.sqrt([F)[F": 104,
    "vecxt.floatarrays$.sumSIMD([F)F": 81,
    "vecxt.floatarrays$.tan$bang([F)V": 93,
    "vecxt.floatarrays$.tan([F)[F": 104,
    "vecxt.floatarrays$.tanh$bang([F)V": 93,
    "vecxt.floatarrays$.tanh([F)[F": 104,
    "vecxt.floatarrays$.unary_$minus([F)[F": 104,
    "vecxt.intarrays$.$bang$colon$eq([II)[Z": 94,
    "vecxt.intarrays$.$bang$colon$eq([I[I)[Z": 106,
    "vecxt.intarrays$.$eq$colon$eq([II)[Z": 94,
    "vecxt.intarrays$.$eq$colon$eq([I[I)[Z": 106,
    "vecxt.intarrays$.$greater$eq([II)[Z": 94,
    "vecxt.intarrays$.$greater$eq([I[I)[Z": 106,
    "vecxt.intarrays$.$greater([II)[Z": 94,
    "vecxt.intarrays$.$greater([I[I)[Z": 106,
    "vecxt.intarrays$.$less$eq([II)[Z": 94,
    "vecxt.intarrays$.$less$eq([I[I)[Z": 106,
    "vecxt.intarrays$.$less([II)[Z": 94,
    "vecxt.intarrays$.$less([I[I)[Z": 106,
    "vecxt.intarrays$.$minus$eq([II)V": 67,
    "vecxt.intarrays$.$minus$eq([I[I)V": 84,
    "vecxt.intarrays$.$minus([II)[I": 16,
    "vecxt.intarrays$.$minus([I[I)[I": 24,
    "vecxt.intarrays$.$plus$eq([I[I)V": 84,
    "vecxt.intarrays$.$plus([I[I)[I": 24,
    "vecxt.intarrays$.countsToIdx([I)[I": 70,
    "vecxt.intarrays$.dot([I[I)I": 108,
    "vecxt.intarrays$.increments([I)[I": 121,
    "vecxt.intarrays$.maxSIMD([I)I": 82,
    "vecxt.intarrays$.mean([I)D": 11,
    "vecxt.intarrays$.meanAndVariance([I)Lvecxt/MeanAndVariance;": 9,
    "vecxt.intarrays$.minSIMD([I)I": 82,
    "vecxt.intarrays$.std([I)D": 9,
    "vecxt.intarrays$.std([ILvecxt/VarianceMode;)D": 10,
    "vecxt.intarrays$.stdDev([I)D": 9,
    "vecxt.intarrays$.sumSIMD([I)I": 81,
    "vecxt.intarrays$.variance([I)D": 9,
    "vecxt.intarrays$.variance([ILvecxt/VarianceMode;)D": 10,
    "vecxt.intarrays.gt([II)[Z": 9,
    "vecxt.intarrays.gt([I[I)[Z": 9,
    "vecxt.intarrays.gte([II)[Z": 9,
    "vecxt.intarrays.gte([I[I)[Z": 9,
    "vecxt.intarrays.lt([II)[Z": 9,
    "vecxt.intarrays.lt([I[I)[Z": 9,
    "vecxt.intarrays.lte([II)[Z": 9,
    "vecxt.intarrays.lte([I[I)[Z": 9,
    "vecxt.intarrays.meanAndVariance([ILvecxt/VarianceMode;)Lvecxt/MeanAndVariance;": 9,
    "vecxt.intarrays.stdDev([ILvecxt/VarianceMode;)D": 9,
    "vecxt.matrix$Layout.linearIndex(II)I": 19,
    "vecxt.matrix$Matrix.colStride()I": 8,
    "vecxt.matrix$Matrix.cols()I": 8,
    "vecxt.matrix$Matrix.hasSimpleContiguousMemoryLayout()Z": 8,
    "vecxt.matrix$Matrix.isDenseColMajor()Z": 8,
    "vecxt.matrix$Matrix.isDenseRowMajor()Z": 8,
    "vecxt.matrix$Matrix.numel()I": 8,
    "vecxt.matrix$Matrix.offset()I": 8,
    "vecxt.matrix$Matrix.rowStride()I": 8,
    "vecxt.matrix$Matrix.rows()I": 8,
    "vecxt.ndarray$.mkNDArray(Ljava/lang/Object;[I[II)Lvecxt/ndarray$NDArray;": 13,
    "vecxt.ndarray.shapeArray(Lvecxt/ndarray$NDArray;)[I": 8,
    "vecxt.ndarrayOps.expandDims(Lvecxt/ndarray$NDArray;I)Lvecxt/ndarray$NDArray;": 9
  }
}

Simon Parten and others added 2 commits August 4, 2026 15:32
…y with D1

The reason it sat commented out since #110 turns out to be structural rather than
incidental: it had no canary. Every assertion in it reads "still zero with EA
off", which is precisely what a run with EA still *on* produces — so the scope
passed whether or not -XX:-DoEscapeAnalysis reached the JVM. That is not a check,
and switching it off was the right call at the time.

The canary was available for free and nobody had noticed. `variance(mode)` is the
one kernel in D1Suite whose zero comes from escape analysis rather than from
intrinsification: it reads one field out of the MeanAndVariance that
meanAndVarianceTwoPass returns and discards the other, so the object is dead and
EA removes it. That object is an ordinary final class, not a Vector, so nothing
intrinsifies it away — with EA off it must reach the heap. Asserting that it *does*
allocate proves the flag took effect, and a failure there says every other
assertion in the scope is passing for the wrong reason.

That is also why it is the canary rather than a 29th kernel assertion: asserting
"≤ 8 bytes/op with EA off" for an EA-dependent kernel would be asserting the
opposite of what the flag does. The distinction the whole scope rests on is
intrinsification-eliminated (survives EA-off) versus EA-eliminated (does not), and
variance is the only member of the second category.

Coverage 14 -> 28 kernels plus the canary, so every @AllocFree method is now
asserted in both scopes. The drift is worth noting as a hazard in itself: this
suite sat at fourteen while D1Suite grew to twenty-nine, and the lists cannot be
factored into a shared collection because the assertion helpers must stay inline
for each call site to get a monomorphic measurement loop.

Docs corrected in two places where I had overstated this scope's reach. It relies
on EA being what rescues a software-path kernel, and that is not always so — D6's
canary is a software-path kernel that allocates with EA *enabled*, so D1Suite
catches that one unaided. What this scope adds is the narrower set of fallbacks
whose objects happen not to escape. Cheap and worth having; not the full substitute
for confirming intrinsification that "what replaces D2" implied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@Quafadas
Quafadas merged commit 41f359a into main Aug 4, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant