What's need to update later: - Best practices related - add `MEGATRON_LOGGING_LEVEL=20` if you want to check the GTP prefetch chain - `PYTORCH_CUDA_ALLOC_CONF` `expandable_segments:True,graph_capture_record_stream_reuse:True,pinned_max_round_threshold_mb:128` - Don't set `CUDA_DEVICE_MAX_CONNECTIONS` for Blackwell+ - Latest perf / scalability on nemotron-next - Update supporting matrix for now - Others...
What's need to update later:
Best practices related
MEGATRON_LOGGING_LEVEL=20if you want to check the GTP prefetch chainPYTORCH_CUDA_ALLOC_CONFexpandable_segments:True,graph_capture_record_stream_reuse:True,pinned_max_round_threshold_mb:128CUDA_DEVICE_MAX_CONNECTIONSfor Blackwell+Latest perf / scalability on nemotron-next
Update supporting matrix for now
Others...