A rail-optimized architecture is intentionally designed so that communication among corresponding GPU ranks within a single scalable unit can remain on the rail or leaf switches. This provides single-hop forwarding and avoids unnecessary traversal through spine switches, reducing latency for collective communication.
Spines become necessary when the deployment expands beyond one scalable unit. Cisco explicitly documents that multiple scalable units can be interconnected through leaf-and-spine switching. Traffic among GPUs inside one scalable unit benefits from rail-optimized single-hop communication, whereas traffic between GPUs located in different scalable units traverses the spine layer.
Therefore, option C precisely describes when spine switches provide architectural value: scaling the AI fabric beyond the boundaries of one scalable unit .
Option A is inaccurate because the rail design already provides the required local GPU bandwidth within the scalable unit; adding spines is not primarily intended to increase local rail bandwidth. Option B mischaracterizes the principal design purpose. Although additional topology can contribute to resiliency, spines are introduced fundamentally for interconnecting scalable units and enabling scale-out. Option D is also incorrect because increased buffering is not the architectural reason for introducing a spine tier.
Study Guide Reference: AI Infrastructure Components and Architecture — rail-optimized topology, scalable units, nonblocking network design, latency, and horizontal scaling.
===============