Teaching Robots to Build the Hardware That Powers AI: Inside NVIDIA’s GB300 Tray Assembly Effort
Key points
- NVIDIA's Seattle Robotics Lab and the Isaac engineering team worked on two GB300 tester-tray tasks, busbar assembly and multi-connector insertion, chosen with NVIDIA Operations and contract manufacturer Foxconn 1.
- Manufacturers asked for a 99.5% success rate and a cycle time of no more than 124 seconds, twice that of skilled workers, with no collisions with the tray allowed.
- A classical pipeline of separate perception, planning and control modules, not end-to-end learning, achieved success rates above 95% on busbar assembly, though screwdriving pushed the full cycle to 160 seconds.
- The team's approach was to use the right tool at the right time: classical baselines first, then imitation learning, reinforcement learning and vision-language-action models only where those baselines fell short.
Why GB300 Trays Are Hard
The GB300 superchip drives AI training and inference for modern foundation models, yet putting its trays together still depends on skilled human labor in factories worldwide. NVIDIA’s Seattle Robotics Lab, working with the Isaac engineering team, asked whether robots could take on that work. The answer was yes, but the lab describes it as one of the hardest problems it has tackled in nine years of research 1.
The team focused on tester trays, which verify GB300 compute modules before shipping. With input from NVIDIA Operations and Foxconn, they chose two tasks. In busbar assembly, a long, heavy metal bar must be grasped, carried to the tray, inserted and screwed down at 16 locations, with fixtures and clamps added and removed along the way. In multi-connector insertion, two large and two small cable-mounted connectors must be lifted out of the tray and pushed into tight-clearance sockets.
Moravec's Paradox on the Factory Floor
Human workers do these jobs fluidly and with little mental effort, which is a textbook case of Moravec’s paradox: what feels easy for people is often brutally hard for machines. Two challenges stand out. The first is uncertainty. Part shapes, appearance and starting poses vary, and the connectors sit on cables that deform, differ from unit to unit and change over time through fatigue.
Factories usually tame uncertainty with fixed automation and fixtures that hold parts to submillimeter accuracy. Rapid design cycles and low volumes make that impractical here, so the robots must adapt on their own. The second challenge is performance. Research often treats 80% to 90% success as solved and tests only 10 to 20 trials, but the team set its bar with NVIDIA Operations and Foxconn so that solved truly means solved. The lab says it is rapidly approaching those thresholds.
The Right Tool at the Right Time
The team’s stated strategy was to match the tool to the subproblem. It implemented classical baselines first, and only when they fell short did it turn to imitation learning, simulation-based reinforcement learning with sim-to-real transfer, real-world reinforcement learning and vision-language-action models. The lab argues this sequence clarifies each approach’s strengths and weaknesses and justifies added complexity, a contrast to a research culture that prizes novelty over well-tuned baselines.
Busbar Assembly: Classics Win
The busbar task has five steps: insert a limit fixture, insert the busbar with its clamp, drive 16 screws, release the clamp, and remove the fixture. Manufacturers asked for a 99.5% success rate and a cycle of at most 124 seconds, with zero unintended contact with the tray, since even minor damage could mean scrapping a system.
The researchers expected to pivot to end-to-end learning, but a modular pipeline of perception, planning and control proved effective enough that no pivot was needed. It was also easier to debug, tune and spread across multiple arms. For perception, they used NVIDIA FoundationPose for 6D pose estimation, which sometimes erred by 90 or 180 degrees when objects were poorly visible. Multi-view input with confidence-based selection reduced those errors, and the team later moved to a specialist model called DOPER. Planning combined a fast free-space waypoint planner with simple alignment motions such as Lissajous curves. Control relied on a high-performance impedance controller with automatic damping design and inertial compensation, which the team sees as a powerful but underused tool in robot learning.
Work was split across three arms: two Flexiv Rizon 4S units, one holding a camera and one handling the fixture, busbar and clamp, plus a Universal Robots UR10e with an OnRobot screwdriver. The clamp was lightly modified so one arm could unlatch it. Success rates exceeded 95%, with remaining failures mostly from grasp errors and in-hand slippage. The fixture and busbar steps met their timing, but screwdriving did not, leaving a 160-second cycle against the 124-second target.
Companies mentioned: Nvidia (NVDA $237.47 ▼0.7%) • Foxconn
Primary sources
The Machines that Make the Machines (developer.nvidia.com) – An NVIDIA developer blog post from the Seattle Robotics Lab and the Isaac engineering team on teaching robots to assemble GB300 tester trays. It explains why the work matters, then describes the two chosen tasks, busbar assembly and multi-connector insertion, selected with NVIDIA Operations and Foxconn. It frames the difficulty around uncertainty in parts and cables and strict performance requirements, including a 99.5% success rate and a 124-second cycle. The team describes a strategy of trying classical methods first, then learning-based approaches only where needed. For busbar assembly, a modular perception, planning and control pipeline using FoundationPose (later DOPER), waypoint planning and impedance control reached over 95% success across three robot arms, with screwdriving making the cycle 160 seconds instead of 124. The provided text cuts off before the multi-connector insertion results. Published on developer.nvidia.com.

Powered by News Ranker