L
LambdaSan Francisco Bay Area

Engineering Manager

On-siteFull Time$330k - $440k per yearPosted 10 days ago

About the role

Who you are

  • We value diverse backgrounds, experiences, and skills, and we're excited to hear from candidates who bring a unique perspective. If you don't exactly meet this description but believe you may be a good fit, please still apply and help us understand your readiness for this role. Your application is not a waste of our time
  • Have 3+ years leading or managing engineers, in AI/ML infrastructure or another large-scale compute environment
  • Have owned production systems with real SLAs, and can balance keeping things running against long-term, high-impact work — paying down toil and technical debt along the way
  • Work confidently in Linux and can debug across the OS, hardware, and networking layers
  • Can lead technical design on medium-to-large efforts: take an ambiguous problem, write the doc, drive alignment across teams, and ship
  • Work well under deadlines and structured project plans, and can tactfully negotiate changes to timelines when reality demands it
  • Collaborate effectively with peer engineering managers on efforts that cut across deployment and operations
  • Build high-performing teams deliberately — through hiring, upskilling, planned skills redundancy, performance management, and clear expectations
  • Have excellent problem-solving and troubleshooting instincts
  • Are excited about working at the intersection of hardware, software, and physical datacenter builds
  • Leave systems, and the teammates around you, better than you found them
  • Depth in any one of these is a strong signal, and helps us place you on the right team. Nobody has all of them
  • Linux systems administration, TCP/IP networking, automation, and scripting
  • Bare metal provisioning and lifecycle management — PXE, Redfish, IPMI, BMC, DHCP, DNS
  • Strong coding ability in at least one language, plus comfort with APIs, distributed systems, and automation pipelines
  • The technologies underpinning our cloud business: GPU acceleration, virtualization, cloud computing
  • Datacenter physical infrastructure: racks, switches, InfiniBand fabric, power domains
  • Network source-of-truth or DCIM tooling (NetBox or similar), and data quality practice at scale
  • Building Linux distributions, or managing OS customization and imaging
  • Incorporating AI-assisted development tools into engineering workflows — code generation, debugging, test development, documentation
  • Customer awareness, empathy, and diplomacy
  • Bachelor's degree or equivalent experience in a technical field What the job involves
  • Fleet Engineering owns the full lifecycle of Lambda's production systems infrastructure — new product introduction, deployment, operation, and reliability of our GPU fleet. We enable the building and running of that infrastructure with speed, ease, and quality. The Fleet Engineering teams:
  • HPC Deployments — Turns bare metal into production-ready capacity: ensure firmware leveling, system burn-in to shake out early failures, and validation of server performance and correctness, through to the InfiniBand fabric and GPU clusters
  • Fleet Reliability — Day-2 operations across the fleet. Keeps systems healthy and keeps as much of the fleet in service for as much of its useful life as possible
  • Fleet Orchestration / Data — Owns our production source of truth system. Synchronizes data from upstream systems and holds the line on correctness and quality, because everything automated downstream depends on it
  • Fleet Orchestration / Automation — Owns the workflow orchestration system which people use to safely work on fleet systems for workflows that include: locking hosts, running firmware leveling jobs, OS installs, burn-in and validation, and reporting on work in flight and its results
  • Fleet Foundation — Builds the host enablement tooling systems: OS and ZTP switch provisioning, firmware management, out-of-band access, and power management
  • The work is highly cross-functional, carries executive visibility, and has a direct impact on Lambda and our customers. Fleet Engineering is at the forefront of delivering on-time, high-quality GPU capacity while driving efficiency at scale
  • We are hiring multiple Engineering Managers for the following teams: Fleet Reliability, HPC Deployments, Fleet Foundation, Fleet Orchestration / Automation. This is a single application for all of them: you apply once, we get to know you, and we match you to the team where your strengths land best
  • Lead and grow a distributed team of top-talent engineers responsible for the deployment and operation of production systems infrastructure
  • Work cross-functionally to deliver projects and deployments on time, ensuring alignment across stakeholders
  • Identify opportunities for efficiency gains in the tools, processes, and automation that teams across the organization rely on day to day
  • Give stakeholders clear visibility into project progress, risks, and outcomes
  • Participate in qualification efforts for new technologies entering our production deployments
  • Drive outcomes by man
IT Services And IT ConsultingSeries EArtificial IntelligenceMachine Learning