C
Crusoe Energy SystemsSan Francisco, California, USA

Director of Engineering (Flex Compute)

On-siteFull Time$285k - $335k per yearPosted 3 days ago

About the role

  • Own the curtailment orchestration and decision layer: grid-signal ingestion, staged shedding, per-SKU power capping without ever silently dropping a paid workload
  • Land a utility-validated pilot: fast ramp to setpoint, tight accuracy, high-fidelity telemetry, and the test harness that proves it before we commit to a utility
  • Deliver dynamic power management for oversubscription — more GPUs per megawatt with no customer-visible impact
  • Build fast, workload-aware GPU power estimation, validated against fleet telemetry : the engine behind shed forecasting, oversubscription admission, and power planning for new silicon
  • Design for safety: authenticated signal ingress, bounded blast radius, fail-safe defaults, manual backstops
  • Own graceful ride-through of power-loss and grid-stress events, integrated with on-site battery and generation backstops
  • Partner deeply with Data Center Engineering (mechanical, electrical, controls) and Energy teams on interconnection commitments, curtailment program design, and BESS/generation integration
  • Integrate with our cloud control plane across Kubernetes and Slurm fleets : one control plane, not two
  • Make build-vs-leverage calls across vendor power-management stacks and grid integration layers
  • Hire and lead the team from the ground up, keeping headcount sublinear to fleet growth through automation Benefits
  • Health & wellbeing: Comprehensive health benefits designed to support your overall wellness
  • Time away: Paid time off for vacations, family bonding, and unexpected needs
  • 401(k) match: Build your financial future with our 401(k) matching program
  • Mental wellness: Resources and support for your emotional wellbeing and navigating life’s challenges
  • Working knowledge of data-center power systems: utility interconnection, switchgear/UPS/BESS, rack/PDU distribution, power telemetry, and GPU power management
  • 12+ years in software engineering, including 5+ leading engineering teams, ideally taking a system from 0→1 to production scale
  • A track record with safety-critical or physically-actuating systems, where a misfire has a bounded, designed-for worst case
  • Experience with energy markets or grid programs demand response, curtailable-load tariffs, ISO/RTO market signals (e.g., ERCOT, PJM, CAISO) or demonstrated speed to fluency in them
  • Deep experience with distributed control planes, orchestration, or fleet automation (e.g., Temporal, Kubernetes, Slurm)
  • A record of hiring senior engineers and running healthy operations for systems that must never fail silently
  • Strong build-vs-buy judgment and comfort deciding with incomplete data
  • Prior work in the energy sector: power generation, storage, utility software, or grid-scale controls
  • GPU cluster operations or AI cloud infrastructure experience
  • Background in GPU power/performance modeling, DVFS, or turning research prototypes into production systems
  • Experience with checkpoint/preemption for large training jobs or power-aware scheduling
Series E