Lead Solutions Architect, GPU Cluster Solutions
GPU/AI Infrastructure startup
Location: United States (Remote). Travel required.
Up to $240K base, plus bonus and equity
As Lead Architect, GPU Cluster Solutions, you’ll own the end-to-end technical design of GPU cluster deployments for a fast-scaling GPU cloud and AI infrastructure business. You’ll translate client requirements and NVIDIA Reference Architecture into buildable, supportable designs spanning compute, storage, networking, software, and spares strategy, for both prospective and signed engagements.
We’re looking for a hands-on infrastructure architect who can move from a client’s technical requirements to a complete bill of design without losing sight of the real-world constraints that make or break a deployment: site power and cooling, hardware lead times, and the SLA commitments the design has to protect. You’ll be the technical backbone of each engagement, making sure what gets promised can actually be built and stood up.
This is a rare opportunity to shape how large-scale GPU clusters get designed and delivered at a defining moment for the company. If you’re equally comfortable selecting an InfiniBand versus RoCE fabric, formulating a hot/cold sparing plan against a contracted SLA, and adapting a reference design to fit a specific site’s power envelope, we’d love to speak with you.
What you’ll do & achieve
- Own the end-to-end technical design of GPU cluster deployments, from client requirements through NVIDIA Reference Architecture compliance to sparing strategy.
- Design GPU cluster configurations across compute, storage, and networking against NVIDIA Reference Architecture (HGX, NVL72) for each signed engagement.
- Translate client technical requirements into a complete bill of design, covering all necessary compute, storage, and networking components.
- Design network topology and fabric selection, including InfiniBand, RoCE, and high-speed Ethernet options, matched to each client’s workload and performance requirements.
- Incorporate internet, VPN, and firewall connectivity into cluster designs, and specify dedicated point-to-point requirements where needed, including protected optical circuits.
- Formulate and own the hot/cold sparing plan for each deployment to meet contracted SLA commitments, adjusting for data center power and cooling parameters and hardware lead-time constraints.
- Adapt cluster designs to fit site-specific power, cooling, and space constraints, working with the data center operations team.
- Align design decisions with realistic hardware delivery timing in partnership with the supply chain team.
- Partner with the deployments and program management team to ensure designs translate cleanly into buildable, trackable project plans.
- Support acceptance test design and criteria definition, ensuring test procedures validate the as-designed architecture.
Who you are
- 7+ years in solutions architecture, network engineering, or systems engineering supporting GPU, HPC, or large-scale compute infrastructure.
- Deep working knowledge of NVIDIA Reference Architecture (HGX, NVL72) and GPU cluster design principles.
- Hands-on experience with InfiniBand, RoCE, and high-speed Ethernet fabric design.
- Experience designing sparing and spares strategies for mission-critical infrastructure.
- Experience incorporating firewall, VPN, and dedicated circuit requirements (e.g. protected optical) into network designs.
- Experience with high-speed shared storage solutions (e.g. Weka, Vast, DDN).
- Experience designing clusters for large enterprise clients or neoclouds, not just internal infrastructure, is a strong advantage.
- Familiarity with NVIDIA NCP program requirements and certification processes, and with capacity or sparing modelling tools, is a plus.
.
