W0710 07:23:54.764000 140531942191552 torch/distributed/run.py:793] W0710 07:23:54.764000 140531942191552 torch/distributed/run.py:793] ***************************************** W0710 07:23:54.764000 140531942191552 torch/distributed/run.py:793] Setting OMP_NUM_THREADS environment variable for each process to be 1 in default, to avoid your system being overloaded, please further tune the variable for optimal performance in your application as needed. W0710 07:23:54.764000 140531942191552 torch/distributed/run.py:793] ***************************************** sg-fuxc2-260708-12652-default0-0:30966:30966 [0] NCCL INFO Bootstrap : Using eth0:10.244.49.68<0> sg-fuxc2-260708-12652-default0-0:30966:30966 [0] NCCL INFO cudaDriverVersion 13000 sg-fuxc2-260708-12652-default0-0:30966:30966 [0] NCCL INFO NCCL version 2.22.3+cuda12.6 sg-fuxc2-260708-12652-default0-0:30967:30967 [0] NCCL INFO cudaDriverVersion 13000 sg-fuxc2-260708-12652-default0-0:30967:30967 [0] NCCL INFO Bootstrap : Using eth0:10.244.49.68<0> sg-fuxc2-260708-12652-default0-0:30967:30967 [0] NCCL INFO NCCL version 2.22.3+cuda12.6 sg-fuxc2-260708-12652-default0-0:30966:30982 [0] NCCL INFO Plugin Path : /opt/hpcx/nccl_rdma_sharp_plugin/lib/libnccl-net.so sg-fuxc2-260708-12652-default0-0:30966:30982 [0] NCCL INFO P2P plugin IBext_v8 sg-fuxc2-260708-12652-default0-0:30966:30982 [0] NCCL INFO NET/IB : No device found. sg-fuxc2-260708-12652-default0-0:30966:30982 [0] NCCL INFO NET/IB : No device found. sg-fuxc2-260708-12652-default0-0:30966:30982 [0] NCCL INFO NET/Socket : Using [0]eth0:10.244.49.68<0> sg-fuxc2-260708-12652-default0-0:30966:30982 [0] NCCL INFO Using network Socket sg-fuxc2-260708-12652-default0-0:30967:30983 [0] NCCL INFO Plugin Path : /opt/hpcx/nccl_rdma_sharp_plugin/lib/libnccl-net.so sg-fuxc2-260708-12652-default0-0:30967:30983 [0] NCCL INFO P2P plugin IBext_v8 sg-fuxc2-260708-12652-default0-0:30967:30983 [0] NCCL INFO NET/IB : No device found. sg-fuxc2-260708-12652-default0-0:30967:30983 [0] NCCL INFO NET/IB : No device found. sg-fuxc2-260708-12652-default0-0:30967:30983 [0] NCCL INFO NET/Socket : Using [0]eth0:10.244.49.68<0> sg-fuxc2-260708-12652-default0-0:30967:30983 [0] NCCL INFO Using network Socket sg-fuxc2-260708-12652-default0-0:30967:30983 [0] NCCL INFO ncclCommInitRank comm 0x558969e41c40 rank 1 nranks 2 cudaDev 0 nvmlDev 0 busId 1a000 commId 0xd46a8c3fdc5326f9 - Init START sg-fuxc2-260708-12652-default0-0:30966:30982 [0] NCCL INFO ncclCommInitRank comm 0x556836373950 rank 0 nranks 2 cudaDev 0 nvmlDev 0 busId 1a000 commId 0xd46a8c3fdc5326f9 - Init START sg-fuxc2-260708-12652-default0-0:30967:30983 [0] init.cc:738 NCCL WARN Duplicate GPU detected : rank 1 and rank 0 both on CUDA device 1a000 sg-fuxc2-260708-12652-default0-0:30966:30982 [0] init.cc:738 NCCL WARN Duplicate GPU detected : rank 0 and rank 1 both on CUDA device 1a000 sg-fuxc2-260708-12652-default0-0:30967:30983 [0] NCCL INFO init.cc:1408 -> 5 sg-fuxc2-260708-12652-default0-0:30966:30982 [0] NCCL INFO init.cc:1408 -> 5 sg-fuxc2-260708-12652-default0-0:30967:30983 [0] NCCL INFO group.cc:70 -> 5 [Async thread] sg-fuxc2-260708-12652-default0-0:30966:30982 [0] NCCL INFO group.cc:70 -> 5 [Async thread] sg-fuxc2-260708-12652-default0-0:30967:30967 [0] NCCL INFO group.cc:420 -> 5 sg-fuxc2-260708-12652-default0-0:30966:30966 [0] NCCL INFO group.cc:420 -> 5 sg-fuxc2-260708-12652-default0-0:30967:30967 [0] NCCL INFO group.cc:546 -> 5 sg-fuxc2-260708-12652-default0-0:30966:30966 [0] NCCL INFO group.cc:546 -> 5 sg-fuxc2-260708-12652-default0-0:30967:30967 [0] NCCL INFO init.cc:1798 -> 5 sg-fuxc2-260708-12652-default0-0:30966:30966 [0] NCCL INFO init.cc:1798 -> 5 [rank1]: Traceback (most recent call last): [rank1]: File "/root/nccl-learning/scripts/28_ch04_process_groups.py", line 122, in [rank1]: main() [rank1]: File "/root/nccl-learning/scripts/28_ch04_process_groups.py", line 59, in main [rank1]: dist.all_reduce(value) [rank1]: File "/usr/local/lib/python3.10/dist-packages/torch/distributed/c10d_logger.py", line 83, in wrapper [rank1]: return func(*args, **kwargs) [rank1]: File "/usr/local/lib/python3.10/dist-packages/torch/distributed/distributed_c10d.py", line 2486, in all_reduce [rank1]: work = group.allreduce([tensor], opts) [rank1]: torch.distributed.DistBackendError: NCCL error in: /opt/pytorch/pytorch/torch/csrc/distributed/c10d/NCCLUtils.hpp:311, invalid usage (run with NCCL_DEBUG=WARN for details), NCCL version 2.22.3 [rank1]: ncclInvalidUsage: This usually reflects invalid usage of NCCL library. [rank1]: Last error: [rank1]: Duplicate GPU detected : rank 1 and rank 0 both on CUDA device 1a000 [rank0]: Traceback (most recent call last): [rank0]: File "/root/nccl-learning/scripts/28_ch04_process_groups.py", line 122, in [rank0]: main() [rank0]: File "/root/nccl-learning/scripts/28_ch04_process_groups.py", line 59, in main [rank0]: dist.all_reduce(value) [rank0]: File "/usr/local/lib/python3.10/dist-packages/torch/distributed/c10d_logger.py", line 83, in wrapper [rank0]: return func(*args, **kwargs) [rank0]: File "/usr/local/lib/python3.10/dist-packages/torch/distributed/distributed_c10d.py", line 2486, in all_reduce [rank0]: work = group.allreduce([tensor], opts) [rank0]: torch.distributed.DistBackendError: NCCL error in: /opt/pytorch/pytorch/torch/csrc/distributed/c10d/NCCLUtils.hpp:311, invalid usage (run with NCCL_DEBUG=WARN for details), NCCL version 2.22.3 [rank0]: ncclInvalidUsage: This usually reflects invalid usage of NCCL library. [rank0]: Last error: [rank0]: Duplicate GPU detected : rank 0 and rank 1 both on CUDA device 1a000 [rank1]:[W710 07:23:57.614660331 ProcessGroupNCCL.cpp:1207] Warning: WARNING: process group has NOT been destroyed before we destruct ProcessGroupNCCL. On normal program exit, the application should call destroy_process_group to ensure that any pending NCCL operations have finished in this process. In rare cases this process can exit before this point and block the progress of another member of the process group. This constraint has always been present, but this warning has only been added since PyTorch 2.4 (function operator()) [rank0]:[W710 07:23:57.630011267 ProcessGroupNCCL.cpp:1207] Warning: WARNING: process group has NOT been destroyed before we destruct ProcessGroupNCCL. On normal program exit, the application should call destroy_process_group to ensure that any pending NCCL operations have finished in this process. In rare cases this process can exit before this point and block the progress of another member of the process group. This constraint has always been present, but this warning has only been added since PyTorch 2.4 (function operator()) E0710 07:23:57.302000 140531942191552 torch/distributed/elastic/multiprocessing/api.py:863] failed (exitcode: 1) local_rank: 0 (pid: 30966) of binary: /usr/bin/python Traceback (most recent call last): File "/usr/local/bin/torchrun", line 33, in sys.exit(load_entry_point('torch==2.5.0a0+872d972e41.nv24.8', 'console_scripts', 'torchrun')()) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 355, in wrapper return f(*args, **kwargs) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/run.py", line 919, in main run(args) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/run.py", line 910, in run elastic_launch( File "/usr/local/lib/python3.10/dist-packages/torch/distributed/launcher/api.py", line 138, in __call__ return launch_agent(self._config, self._entrypoint, list(args)) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/launcher/api.py", line 269, in launch_agent raise ChildFailedError( torch.distributed.elastic.multiprocessing.errors.ChildFailedError: ============================================================ /root/nccl-learning/scripts/28_ch04_process_groups.py FAILED ------------------------------------------------------------ Failures: [1]: time : 2026-07-10_07:23:57 host : sg-fuxc2-260708-12652-default0-0.sg-fuxc2-260708-12652.crater-workspace.svc.cluster.local rank : 1 (local_rank: 1) exitcode : 1 (pid: 30967) error_file: traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html ------------------------------------------------------------ Root Cause (first observed failure): [0]: time : 2026-07-10_07:23:57 host : sg-fuxc2-260708-12652-default0-0.sg-fuxc2-260708-12652.crater-workspace.svc.cluster.local rank : 0 (local_rank: 0) exitcode : 1 (pid: 30966) error_file: traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html ============================================================