Job Opportunity: Full-Time MLOps/ML Systems Engineer - Building Large-Scale LLM Infra
Job Opportunity: Full-Time MLOps/ML Systems Engineer - Building Large-Scale Language Models
Overview
A full-time, W-2 position available for a MLOps/ML Systems Engineer responsible for building the infrastructure layer for large-scale Language Models (LLMs).
The ideal candidate has 2+ years of experience in ML systems and infrastructure, with experience in production PyTorch or JAX on A100, H100, or TPUs. No concurrent engagements.
Key Responsibilities
- Building Infra Layer: Construct the infrastructure layer for large-scale Language Models (LLMs), focusing on GPU kernels, distributed training, or high-throughput inference serving (vLLM/SGLang/TensorRT-LLM).
- Developing Infrastructure: Design and implement the necessary infrastructure components, including but not limited to data pipelines, model serving, and training.
- Collaboration with Teams: Collaborate with cross-functional teams, including data scientists, engineers, and product managers, to ensure seamless integration and efficient performance.
- Performance Optimization: Optimize the infrastructure to achieve high performance and scalability, ensuring that the models run efficiently and reliably.
- Documentation and Communication: Maintain detailed documentation and provide clear communication to stakeholders, including the ability to present complex infrastructure concepts and solutions.
Requirements
- 2+ years of experience in ML systems and infrastructure with experience in production PyTorch or JAX on A100, H100, or TPUs.
- No other concurrent engagements.
Further Details
Interested candidates can learn more about this opportunity on LinkedIn at Opportunity for MLOps Engineer - LLM Systems.