Senior Solutions Architect (Post-Sales)
Venture-backed AI Infrastructure Scale-up
Location: United Kingdom (Remote)
£165k OTE (£150k base + 10% bonus) + equity & benefits
As Senior Solutions Architect, you'll be the technical partner enterprise customers rely on to deploy, run, and scale demanding AI/ML workloads on a modern GPU-accelerated platform. This is a hands-on, customer-facing role: you'll own complex deployments from first workshop to production, and turn tangled infrastructure requirements into architecture that actually holds up under load.
Day to day you'll work alongside platform engineering, MLOps, data science, and infrastructure teams, guiding them through onboarding, tuning workload performance, and keeping large GPU clusters reliable and cost-efficient. When production misbehaves, you're the one leading the investigation and getting things back on track.
It's a rare seat at a company operating right at the centre of the AI buildout, working with the frameworks and silicon most engineers only read about. If you're equally at home designing a multi-tenant ML platform, running a proof-of-concept, and troubleshooting distributed training performance at 2am, I'd like to hear from you.
What you'll do & achieve
- Design end-to-end AI/ML platform architectures across inference, training, and data pipelines
- Build reference architectures for GPU cluster deployment, model serving, and multi-tenant ML infrastructure
- Evaluate and recommend inference serving frameworks (vLLM, TGI, Triton, and similar)
- Advise on GPU fabric topology for distributed training (NVLink, InfiniBand, RoCEv2)
- Shape observability strategies across GPU metrics, OTel, eBPF, and cluster telemetry
- Deliver technical presentations, workshops, and proof-of-concept engagements
- Act as the primary technical advisor and escalation point for your enterprise accounts
- Monitor and troubleshoot production: GPU utilisation, workload performance, cluster health, and cost
- Lead root cause analysis and remediation on the hard, cross-team issues
- Feed insight back to Product and Engineering to influence platform capability and roadmap
- Document reference architectures and implementation guides, and mentor others as the team grows
Who you are
- 8+ years in infrastructure, platform, or solutions engineering, with 3+ years focused on AI/ML infrastructure or MLOps
- Deep Kubernetes expertise: cluster lifecycle, workloads, operators, RBAC
- Hands-on with NVIDIA GPU infrastructure (latest-generation accelerators preferred)
- Fluent in distributed training (NCCL, tensor and pipeline parallelism) and LLM inference serving (vLLM, TGI, and similar)
- Familiar with GPU Operator, MIG, SR-IOV, and high-performance network fabrics
- Strong scripting and automation skills (Python, Bash, Go preferred)
- Comfortable across AWS, Azure, or GCP, including networking, IAM, and managed Kubernetes
- Working knowledge of observability tooling (Prometheus, Grafana, OpenTelemetry)
- A credible communicator who can hold their own with engineers and executives alike
- Bonus points for Run:AI or Slurm experience, GPU scheduling and autoscaling, or certifications like CKA, CKAD, or a cloud Solutions Architect credential
