L
LambdaSan Francisco Bay Area
Engineering Manager
On-siteFull Time$330k - $440k per yearPosted 10 days ago
About the role
Who you are
- We value diverse backgrounds, experiences, and skills, and we're excited to hear from candidates who bring a unique perspective. If you don't exactly meet this description but believe you may be a good fit, please still apply and help us understand your readiness for this role. Your application is not a waste of our time
- Have 3+ years leading or managing engineers, in AI/ML infrastructure or another large-scale compute environment
- Have owned production systems with real SLAs, and can balance keeping things running against long-term, high-impact work — paying down toil and technical debt along the way
- Work confidently in Linux and can debug across the OS, hardware, and networking layers
- Can lead technical design on medium-to-large efforts: take an ambiguous problem, write the doc, drive alignment across teams, and ship
- Work well under deadlines and structured project plans, and can tactfully negotiate changes to timelines when reality demands it
- Collaborate effectively with peer engineering managers on efforts that cut across deployment and operations
- Build high-performing teams deliberately — through hiring, upskilling, planned skills redundancy, performance management, and clear expectations
- Have excellent problem-solving and troubleshooting instincts
- Are excited about working at the intersection of hardware, software, and physical datacenter builds
- Leave systems, and the teammates around you, better than you found them
- Depth in any one of these is a strong signal, and helps us place you on the right team. Nobody has all of them
- Linux systems administration, TCP/IP networking, automation, and scripting
- Bare metal provisioning and lifecycle management — PXE, Redfish, IPMI, BMC, DHCP, DNS
- Strong coding ability in at least one language, plus comfort with APIs, distributed systems, and automation pipelines
- The technologies underpinning our cloud business: GPU acceleration, virtualization, cloud computing
- Datacenter physical infrastructure: racks, switches, InfiniBand fabric, power domains
- Network source-of-truth or DCIM tooling (NetBox or similar), and data quality practice at scale
- Building Linux distributions, or managing OS customization and imaging
- Incorporating AI-assisted development tools into engineering workflows — code generation, debugging, test development, documentation
- Customer awareness, empathy, and diplomacy
- Bachelor's degree or equivalent experience in a technical field What the job involves
- Fleet Engineering owns the full lifecycle of Lambda's production systems infrastructure — new product introduction, deployment, operation, and reliability of our GPU fleet. We enable the building and running of that infrastructure with speed, ease, and quality. The Fleet Engineering teams:
- HPC Deployments — Turns bare metal into production-ready capacity: ensure firmware leveling, system burn-in to shake out early failures, and validation of server performance and correctness, through to the InfiniBand fabric and GPU clusters
- Fleet Reliability — Day-2 operations across the fleet. Keeps systems healthy and keeps as much of the fleet in service for as much of its useful life as possible
- Fleet Orchestration / Data — Owns our production source of truth system. Synchronizes data from upstream systems and holds the line on correctness and quality, because everything automated downstream depends on it
- Fleet Orchestration / Automation — Owns the workflow orchestration system which people use to safely work on fleet systems for workflows that include: locking hosts, running firmware leveling jobs, OS installs, burn-in and validation, and reporting on work in flight and its results
- Fleet Foundation — Builds the host enablement tooling systems: OS and ZTP switch provisioning, firmware management, out-of-band access, and power management
- The work is highly cross-functional, carries executive visibility, and has a direct impact on Lambda and our customers. Fleet Engineering is at the forefront of delivering on-time, high-quality GPU capacity while driving efficiency at scale
- We are hiring multiple Engineering Managers for the following teams: Fleet Reliability, HPC Deployments, Fleet Foundation, Fleet Orchestration / Automation. This is a single application for all of them: you apply once, we get to know you, and we match you to the team where your strengths land best
- Lead and grow a distributed team of top-talent engineers responsible for the deployment and operation of production systems infrastructure
- Work cross-functionally to deliver projects and deployments on time, ensuring alignment across stakeholders
- Identify opportunities for efficiency gains in the tools, processes, and automation that teams across the organization rely on day to day
- Give stakeholders clear visibility into project progress, risks, and outcomes
- Participate in qualification efforts for new technologies entering our production deployments
- Drive outcomes by man