B
BasetenSan Francisco, California, United States {{REMOTE}}

Engineering Manager (Forward Deployed Engineering, LLM)

On-siteFull Time$260k - $380k per yearPosted 30 days ago

About the role

  • As an Engineering Manager (Player & Coach), you will lead and mentor a team of Forward Deployed Engineers focused on building, scaling, and optimizing LLM inference workloads for Baseten customers
  • Applying both hands-on technical ownership and managerial leadership, you will guide your team through the processes of designing, deploying, and managing high performance, low latency AI applications on Baseten’s platform
  • FDE at baseten is not a sales function – we are a mix of engineering, product, and customer architects who contribute to the core Baseten codebase, drive large portions of our feature roadmap, and execute on complicated customer engagements
  • You will also partner with product, infrastructure, and other customer engineering teams to ensure that large language models (LLMs) and other generative AI systems deliver best-in-class performance, reliability, and cost efficiency in production environments
  • Lead, mentor, and grow a team of Forward Deployed Engineers, providing guidance on technical direction, project execution, and professional development
  • Set clear goals and ensure timely, high-quality delivery across multiple customer-facing projects involving LLM deployment and inference optimization
  • Collaborate with leadership to align team priorities with company and customer goals, balancing short-term delivery, widely varying customer priorities, and long-term technical initiatives
  • Player-coach – While much of this role will be leading the team, you will also be expected to be a key driver on strategic product initiatives and customer engagements. The best managers derive credibility from being able to be hands-on when needed
  • Develop and maintain software systems and product features using one or more general-purpose programming languages in a production-level environment, with a preference for Python due to its relevance in ML projects
  • Drive customer impact by designing, implementing, and deploying Baseten solutions end-to-end (problem framing → evaluation → production deployment → monitoring). This involves working with customers’ engineering teams at every stage of the customer journey including: sales, implementation, and expansion
  • Deliver with velocity: turn vague objectives into clear specs and well-defined PoCs so we can rapidly ship well-tested services and outcomes for our customers
  • Optimize and enhance AI/ML projects, contributing to the continuous improvement of our technical stack. This includes developing features and PRDs with other engineering and product orgs
  • Own products and customer projects end-to-end, functioning as both an engineer, project manager, and product manager, with a focus on user empathy, project specification, and end-to-end execution Benefits
  • Remote-first work environment. The Baseten team is welcome to work from wherever they want; fully remote, in our San Francisco office, or a mix of both. Today, our team (including our founding team) is spread across the United States, Canada, and Armenia. We provide a $1,000 stipend for you to make your home-office comfortable and productive
  • Regular in-person team summits. We get together as a team three times a year to plan, workshop, and most importantly, get to know each other better
  • Unlimited PTO. We ask that everyone take at least 4 weeks of vacation. And we have a company-wide break between Christmas and New Year’s Day
  • Full healthcare coverage. Medical, dental and vision insurance for you and your family
  • Paid parental leave. 16-weeks fully paid parental leave (adoptive and non-birth parents included) and flexibility with schedules while returning to work
  • Company-sponsored 401(k) for you to contribute to
  • Learning and development budget. We encourage you to take classes, attend conferences, and invest in your craft and we’ll cover expenses to make it happen
  • Strong programming skills in Python, with production experience in building or optimizing ML inference systems
  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or related field
  • Proven experience with LLMs, inference optimization, or serving frameworks (e.g., vLLM, TensorRT, Triton, Hugging Face, Ray Serve)
  • Familiarity with observability, profiling, and cost/performance tradeoffs in production ML systems
  • 4+ years of professional software engineering experience, including 1+ year in a leadership or mentorship capacity
  • Excellent communication and collaboration skills—able to lead cross-functional efforts and drive outcomes in ambiguous, fast-paced environments
  • If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you
  • Experience leading customer-facing engineering teams or working directly with enterprise partners
  • Deep understanding of GPU infrastructure, distributed inference, or model compression techniques
Software DevelopmentSeries CMachine LearningDeveloper Tools