Exam NCA-AIIO Topic 2 Question 34 Discussion

Actual exam question for NVIDIA's NCA-AIIO exam
Question #: 34
Topic #: 2
You are configuring a multi-node AI training environment using NVIDIA GPUs, and your team wants to ensure that the network infrastructure can handle the data transfer between nodes efficiently, especially during distributed training tasks. What is the most critical factor to consider in the network infrastructure to minimize bottlenecks during distributed AI training?

Suggested Answer: A Vote an answer

Implementing InfiniBand with RDMA support is the most critical factor to minimize bottlenecks in distributed AI training. It provides ultra-low latency and high bandwidth (e.g., 200 Gb/s), optimizing GPU-to- GPU data transfers via NCCL. Option B (more Ethernet ports) improves redundancy, not speed. Option C (fewer nodes) limits scalability. Option D (SDN) aids management, not raw performance. NVIDIA's DGX networking guides recommend InfiniBand.

by Fabian at Jul 22, 2026, 02:47 AM

Comments

Chosen Answer:
This is a voting comment (?) , you can switch to a simple comment.
Switch to a voting comment New
Nick name: Submit Cancel
A voting comment increases the vote count for the chosen answer by one.

Upvoting a comment with a selected answer will also increase the vote count towards that answer by one. So if you see a comment that you already agree with, you can upvote it instead of posting a new comment.

0
0
0
10