C
CoreWeaveLivingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA
Engineering Manager (Kubernetes Infrastructure, Bare Metal)
On-siteFull Time$182k - $242k per yearPosted 10 days ago
About the role
- CoreWeave is looking for an Engineering Manager to lead a team building and operating Kubernetes infrastructure on bare metal
- This team sits close to the core of our platform and is responsible for the reliability, scalability, and operational excellence of the systems that power high-performance AI and ML workloads
- You will lead engineers working on cluster lifecycle, platform reliability, infrastructure automation, and the operational systems that make Kubernetes run predictably at scale on dedicated hardware
- This is a hands-on leadership role for someone who can grow engineers, improve execution, and partner deeply with platform, networking, compute, and product teams
- The right person understands what it takes to run Kubernetes in demanding production environments and can turn that understanding into a high-functioning team, clear priorities, and durable engineering systems
- As the Engineering Manager for Kubernetes Infrastructure, you will lead a team responsible for the core infrastructure and operational foundations behind Kubernetes running directly on bare metal
- At CoreWeave, this platform is built for high-performance computing workloads and gives customers direct access to dedicated hardware, without virtualization overhead
- That makes reliability, performance, observability, and safe lifecycle management especially important
- You will be responsible for team execution, technical direction in partnership with senior engineers, and building strong operating mechanisms around delivery, quality, and incident response
- You will help the team scale its impact while mentoring engineers, supporting hiring, and strengthening cross-functional collaboration
- Lead a team of engineers responsible for Kubernetes infrastructure running on bare metal
- Set clear goals, priorities, and execution plans for the team, and ensure reliable delivery against them
- Partner with senior ICs and adjacent teams on the roadmap for cluster lifecycle management, upgrades, reliability, observability, and infrastructure automation
- Improve the team’s operational excellence across incident response, on-call health, root-cause analysis, and service ownership
- Drive engineering best practices for safe change management, testing, rollout quality, and production readiness
- Support the design and operation of platform capabilities for provisioning, patching, upgrades, scaling, and troubleshooting of Kubernetes clusters
- Build strong cross-functional relationships with compute, networking, storage, security, and product stakeholders
- Hire, coach, and develop engineers while creating a high-accountability, high-trust team culture
- Establish and improve mechanisms for planning, prioritization, execution tracking, and continuous improvement
- Help translate complex platform and infrastructure work into clear business and customer value
- In the first 90 days
- Build trust with the team and key partner organizations
- Assess team health, role clarity, roadmap risks, and operational pain points
- Establish or tighten core operating rhythms for planning, execution, and incident follow-up
- Create a clear view of the highest-value reliability and scalability opportunities
- In the first 6 months
- Improve predictability of team execution and service ownership
- Raise the quality bar for change management, rollout safety, and operational readiness
- Strengthen hiring and development plans for the team
- Drive measurable improvements in one or more areas such as cluster reliability, upgrade safety, provisioning speed, or observability
- In the first 12 months
- Build a strong, durable team with clear ownership and healthy operating mechanisms
- Deliver meaningful infrastructure improvements that increase platform reliability, scalability, and maintainability
- Be recognized as a trusted cross-functional leader for Kubernetes infrastructure on bare metal
- Experience leading incident response cultures and driving follow-through on reliability improvements
- Experience managing an infrastructure, platform, or SRE-oriented engineering team
- Strong written and verbal communication, including the ability to explain technical trade-offs and priorities clearly
- Track record of improving team execution, engineering quality, and operational maturity
- Strong technical depth in Kubernetes, distributed systems, and production infrastructure
- Experience operating Kubernetes in complex environments, ideally including bare metal, hybrid, or highly performance-sensitive systems
- Strong partnership skills across engineering, product, and operations functions
- Ability to coach engineers at different levels and create clarity in ambiguous or fast-scaling environments
- Familiarity with cluster lifecycle management, including provisioning, upgrades, node operations, observability, and reliability engineering
- Experience with GPU-heavy, HPC, or ML infrastructure environments
- Experience building internal platform products used by other