C
CoreWeaveLivingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA

Engineering Manager (Kubernetes Infrastructure, Bare Metal)

On-siteFull Time$182k - $242k per yearPosted 10 days ago

About the role

  • CoreWeave is looking for an Engineering Manager to lead a team building and operating Kubernetes infrastructure on bare metal
  • This team sits close to the core of our platform and is responsible for the reliability, scalability, and operational excellence of the systems that power high-performance AI and ML workloads
  • You will lead engineers working on cluster lifecycle, platform reliability, infrastructure automation, and the operational systems that make Kubernetes run predictably at scale on dedicated hardware
  • This is a hands-on leadership role for someone who can grow engineers, improve execution, and partner deeply with platform, networking, compute, and product teams
  • The right person understands what it takes to run Kubernetes in demanding production environments and can turn that understanding into a high-functioning team, clear priorities, and durable engineering systems
  • As the Engineering Manager for Kubernetes Infrastructure, you will lead a team responsible for the core infrastructure and operational foundations behind Kubernetes running directly on bare metal
  • At CoreWeave, this platform is built for high-performance computing workloads and gives customers direct access to dedicated hardware, without virtualization overhead
  • That makes reliability, performance, observability, and safe lifecycle management especially important
  • You will be responsible for team execution, technical direction in partnership with senior engineers, and building strong operating mechanisms around delivery, quality, and incident response
  • You will help the team scale its impact while mentoring engineers, supporting hiring, and strengthening cross-functional collaboration
  • Lead a team of engineers responsible for Kubernetes infrastructure running on bare metal
  • Set clear goals, priorities, and execution plans for the team, and ensure reliable delivery against them
  • Partner with senior ICs and adjacent teams on the roadmap for cluster lifecycle management, upgrades, reliability, observability, and infrastructure automation
  • Improve the team’s operational excellence across incident response, on-call health, root-cause analysis, and service ownership
  • Drive engineering best practices for safe change management, testing, rollout quality, and production readiness
  • Support the design and operation of platform capabilities for provisioning, patching, upgrades, scaling, and troubleshooting of Kubernetes clusters
  • Build strong cross-functional relationships with compute, networking, storage, security, and product stakeholders
  • Hire, coach, and develop engineers while creating a high-accountability, high-trust team culture
  • Establish and improve mechanisms for planning, prioritization, execution tracking, and continuous improvement
  • Help translate complex platform and infrastructure work into clear business and customer value
  • In the first 90 days
  • Build trust with the team and key partner organizations
  • Assess team health, role clarity, roadmap risks, and operational pain points
  • Establish or tighten core operating rhythms for planning, execution, and incident follow-up
  • Create a clear view of the highest-value reliability and scalability opportunities
  • In the first 6 months
  • Improve predictability of team execution and service ownership
  • Raise the quality bar for change management, rollout safety, and operational readiness
  • Strengthen hiring and development plans for the team
  • Drive measurable improvements in one or more areas such as cluster reliability, upgrade safety, provisioning speed, or observability
  • In the first 12 months
  • Build a strong, durable team with clear ownership and healthy operating mechanisms
  • Deliver meaningful infrastructure improvements that increase platform reliability, scalability, and maintainability
  • Be recognized as a trusted cross-functional leader for Kubernetes infrastructure on bare metal
  • Experience leading incident response cultures and driving follow-through on reliability improvements
  • Experience managing an infrastructure, platform, or SRE-oriented engineering team
  • Strong written and verbal communication, including the ability to explain technical trade-offs and priorities clearly
  • Track record of improving team execution, engineering quality, and operational maturity
  • Strong technical depth in Kubernetes, distributed systems, and production infrastructure
  • Experience operating Kubernetes in complex environments, ideally including bare metal, hybrid, or highly performance-sensitive systems
  • Strong partnership skills across engineering, product, and operations functions
  • Ability to coach engineers at different levels and create clarity in ambiguous or fast-scaling environments
  • Familiarity with cluster lifecycle management, including provisioning, upgrades, node operations, observability, and reliability engineering
  • Experience with GPU-heavy, HPC, or ML infrastructure environments
  • Experience building internal platform products used by other
Technology, Information And InternetSeries UnknownCloudKubernetes