Skip to content

PyTorch Full Inductor Unit Tests #5

Description

@naromero77amd

gfx1250 PyTorch Inductor Merged Outcome

Updated: 2026-08-05

Important

This is the latest exact-node merge, not a fully completed targeted rerun.
All seven high-MISS full suites completed, but the targeted phase completed
27 / 1,848 nodes and left 1,821 reruns pending.

Overall Result

  • Discovered nodes: 22,962
  • Explicit PASSED outcomes: 17,580
  • Pass rate (passed + skipped + xfailed) / total: 21,096 / 22,962 = 91.87%
  • Failing outcomes (failed + error + missed) / total: 1,866 / 22,962 = 8.13%
  • Merge rule: the latest terminal result overrides by exact pytest node ID.

Suite Summary

Suite Total Passed Skipped Xfailed Failed Error Timed Out Missed
test_alignment.py 12 12 0 0 0 0 0 0
test_analysis.py 28 28 0 0 0 0 0 0
test_aot_inductor.py 950 455 480 3 7 5 0 0
test_aot_inductor_arrayref.py 314 128 182 3 1 0 0 0
test_aot_inductor_custom_ops.py 35 35 0 0 0 0 0 0
test_aot_inductor_package.py 88 59 25 0 4 0 0 0
test_async_compile.py 8 8 0 0 0 0 0 0
test_augmented_graph_helper.py 20 20 0 0 0 0 0 0
test_auto_chunker.py 9 9 0 0 0 0 0 0
test_auto_functionalize.py 39 38 1 0 0 0 0 0
test_benchmark_fusion.py 16 10 4 0 0 2 0 0
test_benchmarking.py 14 8 0 6 0 0 0 0
test_best_config.py 1 1 0 0 0 0 0 0
test_binary_folding.py 6 5 1 0 0 0 0 0
test_block_analysis.py 10 10 0 0 0 0 0 0
test_cache.py 725 725 0 0 0 0 0 0
test_caching.py 212 212 0 0 0 0 0 0
test_ck_backend.py 34 0 34 0 0 0 0 0
test_codecache.py 257 193 59 1 3 1 0 0
test_codegen_triton.py 1 1 0 0 0 0 0 0
test_collective_autotuning.py 2 0 2 0 0 0 0 0
test_combo_kernels.py 77 70 5 0 0 2 0 0
test_compile.py 10 10 0 0 0 0 0 0
test_compile_subprocess.py 936 868 40 0 8 20 0 0
test_compile_worker.py 16 16 0 0 0 0 0 0
test_compiled_autograd.py 875 865 4 5 1 0 0 0
test_compiled_optimizers.py 682 677 3 0 0 2 0 0
test_config.py 14 14 0 0 0 0 0 0
test_control_deps.py 4 4 0 0 0 0 0 0
test_control_flow.py 741 739 2 0 0 0 0 0
test_cooperative_reductions.py 163 160 0 0 0 3 0 0
test_coordinate_descent_tuner.py 5 5 0 0 0 0 0 0
test_cpp_wrapper_hipify.py 3 3 0 0 0 0 0 0
test_cpu_repro.py 752 747 5 0 0 0 0 0
test_cuda_repro.py 98 78 3 1 3 13 0 0
test_cudacodecache.py 3 0 0 0 3 0 0 0
test_cudagraph_trees.py 189 180 8 0 0 1 0 0
test_cudagraph_trees_expandable_segments.py 158 151 6 0 0 1 0 0
test_custom_lowering.py 6 4 1 0 1 0 0 0
test_custom_op_autotune.py 4 3 0 0 1 0 0 0
test_custom_partitioner_fn.py 1 1 0 0 0 0 0 0
test_custom_post_grad_passes.py 6 6 0 0 0 0 0 0
test_cutedsl_grouped_mm.py 24 0 24 0 0 0 0 0
test_cutedsl_template.py 13 0 13 0 0 0 0 0
test_cutlass_backend.py 181 0 181 0 0 0 0 0
test_cutlass_evt.py 8 0 8 0 0 0 0 0
test_debug_trace.py 3 3 0 0 0 0 0 0
test_decompose_mem_bound_mm.py 37 35 2 0 0 0 0 0
test_dependencies.py 5 5 0 0 0 0 0 0
test_deterministic.py 32 7 0 0 24 1 0 0
test_device_assert.py 8 8 0 0 0 0 0 0
test_distributed_patterns.py 20 20 0 0 0 0 0 0
test_efficient_conv_bn_eval.py 2 1 0 0 0 1 0 0
test_exc_lowering_stack_trace.py 2 2 0 0 0 0 0 0
test_extension_backend.py 1 0 0 0 1 0 0 0
test_external_callables.py 3 2 0 0 0 1 0 0
test_flex_attention.py 786 750 27 1 2 6 0 0
test_flex_decoding.py 557 554 1 2 0 0 0 0
test_flex_flash.py 210 6 204 0 0 0 0 0
test_foreach.py 607 581 26 0 0 0 0 0
test_fp8.py 202 98 55 0 44 5 0 0
test_fused_attention.py 116 114 2 0 0 0 0 0
test_fusion_regions.py 6 6 0 0 0 0 0 0
test_fuzzer.py 11 10 1 0 0 0 0 0
test_fx_fusion.py 4 4 0 0 0 0 0 0
test_fxir_backend.py 75 74 0 0 0 1 0 0
test_gpu_cpp_wrapper.py 297 289 1 0 3 4 0 0
test_gpu_select_algorithm.py 58 58 0 0 0 0 0 0
test_graph_transform_observer.py 1 1 0 0 0 0 0 0
test_group_batch_fusion.py 13 11 2 0 0 0 0 0
test_halide.py 4 0 4 0 0 0 0 0
test_helion_kernels.py 2 0 2 0 0 0 0 0
test_indexing.py 22 22 0 0 0 0 0 0
test_inductor_annotations.py 2 2 0 0 0 0 0 0
test_inductor_freezing.py 48 33 1 0 1 2 0 11
test_inductor_scheduler.py 10 10 0 0 0 0 0 0
test_inductor_utils.py 2 2 0 0 0 0 0 0
test_inplace_padding.py 9 9 0 0 0 0 0 0
test_inplacing_pass.py 23 23 0 0 0 0 0 0
test_kernel_benchmark.py 19 16 0 0 0 3 0 0
test_kernel_optimization.py 1 1 0 0 0 0 0 0
test_lookup_table.py 37 29 4 0 0 0 0 4
test_loop_ordering.py 70 68 0 0 0 2 0 0
test_max_autotune.py 472 182 173 0 17 95 0 5
test_max_autotune_blackwell.py 106 2 104 0 0 0 0 0
test_mem_estimation.py 4 4 0 0 0 0 0 0
test_memory.py 8 8 0 0 0 0 0 0
test_memory_planning.py 4 4 0 0 0 0 0 0
test_metrics.py 6 6 0 0 0 0 0 0
test_minifier.py 14 14 0 0 0 0 0 0
test_minifier_isolate.py 2 1 1 0 0 0 0 0
test_minifier_utils.py 3 3 0 0 0 0 0 0
test_mix_order_reduction.py 485 248 219 0 16 1 0 1
test_mkldnn_pattern_matcher.py 20 16 4 0 0 0 0 0
test_mmdecomp.py 28 28 0 0 0 0 0 0
test_move_constructors_to_gpu.py 7 6 1 0 0 0 0 0
test_mps_basic.py 46 24 0 0 20 2 0 0
test_multi_kernel.py 19 16 2 0 0 1 0 0
test_native_matmul.py 14 2 0 0 4 8 0 0
test_needs_exact_strides.py 2 2 0 0 0 0 0 0
test_nv_universal_gemm.py 23 3 20 0 0 0 0 0
test_online_softmax.py 31 31 0 0 0 0 0 0
test_op_completeness.py 5 4 1 0 0 0 0 0
test_op_dtype_prop.py 581 581 0 0 0 0 0 0
test_ordered_set.py 401 386 15 0 0 0 0 0
test_pad_mm.py 19 8 1 0 0 10 0 0
test_padding.py 55 46 9 0 0 0 0 0
test_pattern_matcher.py 63 56 0 0 0 7 0 0
test_perf.py 66 66 0 0 0 0 0 0
test_profiler.py 8 8 0 0 0 0 0 0
test_provenance_tracing.py 16 16 0 0 0 0 0 0
test_quantization.py 2 2 0 0 0 0 0 0
test_remote_cache.py 3 3 0 0 0 0 0 0
test_scatter_optimization.py 8 8 0 0 0 0 0 0
test_segmented_tree.py 12 12 0 0 0 0 0 0
test_select_algorithm.py 27 14 1 0 2 9 0 1
test_selective_lowering.py 2 2 0 0 0 0 0 0
test_smoke.py 3 3 0 0 0 0 0 0
test_snode_runtime.py 22 22 0 0 0 0 0 0
test_split_cat_fx_aten_passes.py 5 5 0 0 0 0 0 0
test_split_cat_fx_passes.py 11 11 0 0 0 0 0 0
test_static_triton_launcher.py 17 17 0 0 0 0 0 0
test_subgraph_choice.py 2 2 0 0 0 0 0 0
test_template_heuristics_registry.py 7 7 0 0 0 0 0 0
test_torchbind.py 16 16 0 0 0 0 0 0
test_torchinductor.py 1041 962 42 0 3 22 0 12
test_torchinductor_codegen_config_overrides.py 4 4 0 0 0 0 0 0
test_torchinductor_codegen_dynamic_shapes.py 1864 1466 200 195 3 0 0 0
test_torchinductor_dynamic_shapes.py 1933 1808 112 3 8 2 0 0
test_torchinductor_opinfo.py 3674 1555 691 38 0 1390 0 0
test_torchinductor_strided_blocks.py 304 86 194 0 1 1 0 22
test_triton_extension_backend.py 3 3 0 0 0 0 0 0
test_triton_helpers.py 2 2 0 0 0 0 0 0
test_triton_heuristics.py 13 11 1 0 1 0 0 0
test_triton_kernels.py 372 326 44 0 0 2 0 0
test_triton_syntax.py 1 1 0 0 0 0 0 0
test_triton_wrapper.py 2 2 0 0 0 0 0 0
test_unbacked_symints.py 34 32 0 0 0 2 0 0
test_utils.py 11 11 0 0 0 0 0 0
test_xpu_basic.py 4 4 0 0 0 0 0 0
Total 22962 17580 3258 258 182 1628 0 56

Execution and Merge Notes

Continuation coverage

  • Full-suite phase: 6,254 / 6,254 completed.
  • Targeted phase: 27 / 1,848 completed; 1,821 pending.
  • Total continuation executions: 6,281 / 8,102.
  • Merged provenance: 16,681 latest outcomes from ROCm/HIP 7.15.0 and 6,281 from ROCm/HIP 7.15.26305.
  • Pending targeted nodes retain their earlier terminal state in the table.

Stop-condition correction

  • The targeted continuation stopped under the previous monitor after it accumulated 10 distinct SIGABRT-producing nodes.
  • Those were not 10 adjacent executed tests. Replaying the completed result sequence gives a maximum adjacent SIGABRT streak of 2.
  • The monitor has been corrected to stop only after 5 immediately adjacent executed tests trigger SIGABRT; any other completed test resets the streak.

Notable outcomes

  • The previously GPU-loss-associated FlexAttention node test_builtin_score_mods_different_block_size_score_mod6_BLOCK_SIZE3_cuda_float16 passed in the continuation.
  • The prior hard-wedge node test_layer_norm_bwd_with_dynamic_shape_dynamic_dims2 remains MISSED and is still pending in the targeted phase.
  • After the stop, no test/compile processes survived and the gfx1250 tensor correctness smoke test passed.
  • The merged artifacts do not retain comparable per-node timing for every source run, so recorded-time totals are intentionally omitted.

Largest unresolved clusters

  • Errors: test_torchinductor_opinfo.py (1,390) and test_max_autotune.py (95).
  • Failures: test_fp8.py (44), test_deterministic.py (24), and test_mps_basic.py (20).
  • Misses: test_torchinductor_strided_blocks.py (22), test_torchinductor.py (12), and test_inductor_freezing.py (11).

Environment

  • GPU: AMD Radeon Graphics, gfx1250
  • Baseline: PyTorch 2.11.0+rocm7.15.0a20260721, ROCm/HIP 7.15.0, Triton 3.8.0
  • Continuation: PyTorch 2.11.0+rocm10.1.0a20260803, ROCm/HIP 7.15.26305, Triton 3.8.0

Detailed reproduction steps and high-risk test notes are in this issue comment.

The comments below are intermediate checkpoints and can be ignored.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions