Skip to content

ggml-cpu: vectorize the fused RMS-norm kernel (ggml_compute_forward_rms_norm_f32) - #32

Open
codspeed-hq[bot] wants to merge 1 commit into
masterfrom
codspeed-optim-vectorize-the-fused-rms-norm-kernel-ggml-compute-f-1785230254392
Open

ggml-cpu: vectorize the fused RMS-norm kernel (ggml_compute_forward_rms_norm_f32)#32
codspeed-hq[bot] wants to merge 1 commit into
masterfrom
codspeed-optim-vectorize-the-fused-rms-norm-kernel-ggml-compute-f-1785230254392

ggml-cpu: vectorize the fused RMS-norm kernel

c1db5d3
Select commit
Loading
Failed to load commit list.
CodSpeed HQ / CodSpeed Performance Analysis succeeded Jul 28, 2026 in 0s

Performance Gate Passed

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 1 improved benchmark
✅ 27 untouched benchmarks

Performance Changes

Mode Benchmark BASE HEAD Efficiency
Simulation rms_norm 561.2 µs 323.4 µs +73.53%

Tip

Curious why this is faster? Comment @codspeedbot explain why this is faster on this PR, or directly use the CodSpeed MCP with your agent.


Comparing codspeed-optim-vectorize-the-fused-rms-norm-kernel-ggml-compute-f-1785230254392 (c1db5d3) with master (46819c9)

Open in CodSpeed