Job opening

Senior HPC Deployment and Cable Validation :: Santa Clara, CA :: W2

Prudent Technologies and Consulting, Inc.

Santa Clara, CA

Filed under IT Services and IT Consulting

Full job description

Hi,


Role :- Senior HPC Deployment and Cable Validation

Location :- Santa Clara, CA(Onsite)


P O S I T I O N D E S C R I P T I O N

ENGAGEMENT SUMMARY

The Candidate will provide deployment services for HPC and AI cluster infrastructure with an emphasis on


physical build quality, cable validation, rack integration, and cluster readiness. This role is for a senior hands-on


operator who can turn physical deployment work into repeatable, auditable infrastructure quality.

WHAT THIS CANDIDATE WILL BE DOING

Execute and coordinate rack-level deployment of compute, network, and storage infrastructure for AI and


HPC environments.


• Validate rack elevations, port maps, cable maps, power distribution, labeling, and physical connectivity


before cluster handoff.


• Perform detailed cable validation across Ethernet, InfiniBand, management, and storage interconnects.


• Detect and resolve cabling defects including polarity issues, incorrect port destinations, unsupported optics


pairings, breakout errors, damaged media, and inconsistent labeling.


• Support hardware bring-up by validating BIOS baselines, BMC reachability, inventory accuracy, and initial


connectivity tests.


• Partner with network and Linux teams during cluster turn-up to quickly isolate physical-layer versus logical-


layer failures.


• Create deployment checklists, as-built documentation, and signoff criteria for production acceptance.


• Drive structured remediation during expansion, re-cabling, or failed build events.


WHAT WE NEE D TO SEE


• 7+ years in data center deployment, HPC infrastructure installation, or large-scale hardware integration.


• Strong practical knowledge of structured cabling, optics, transceivers, breakout schemes, and rack-level


physical validation.


• Experience reading and validating rack diagrams, patch plans, cable matrices, and topology documentation.


• Strong troubleshooting skill for L1 issues that manifest as network instability, missing hosts, degraded


performance, or failed cluster readiness checks.


• Experience with asset tracking, labeling discipline, and deployment quality control.


• Ability to work across physical infrastructure, server hardware, and network operations teams without


losing detail.

PREFERRED EXPERIENCE

• Experience in GPU cluster deployment or AI factory buildouts.


• Familiarity with automated cable validation or DCIM-integrated deployment workflows.


• Working knowledge of host provisioning and network validation enough to accelerate multi-team turn-up.


Apply on original listing