Applied AI/ML Engineer
Remote (United States)
Job Details
Job Type: Full-Time
Salary: $175,000 - $250,000 per year, plus equity
About the Role
This opportunity is for an Applied AI/ML Engineer responsible for building and delivering production AI products powered by large language models and modern machine learning technologies. You will own the complete development lifecycle, from prototype through production deployment, while optimizing model performance, inference efficiency, scalability, and operational reliability across GPU infrastructure.
This is a highly autonomous engineering role focused on delivering production-ready AI systems rather than research projects. You will work across machine learning, reinforcement learning, inference optimization, GPU infrastructure, and product development while continuously improving quality, throughput, latency, and operational efficiency.
What You'll Do
End-to-End AI Product Delivery
- Own AI-powered features and products from initial prototype through production deployment, including model selection, inference serving, evaluation, optimization, and continuous iteration.
- Deliver production-ready AI software designed for real-world performance rather than research prototypes.
Inference Serving
- Deploy and optimize large language model (LLM) inference using vLLM and SGLang.
- Optimize continuous batching, KV-cache management, quantization, speculative decoding, and multi-model routing to maximize throughput while minimizing latency and inference cost.
Reinforcement Learning & Post-Training
- Build, deploy, and maintain reinforcement learning and post-training pipelines using slime, Megatron-LM, SGLang, prime-rl, and related tooling.
- Design reward functions and verifier systems.
- Manage rollout orchestration, weight synchronization, and long-running training stability.
Evaluation & Continuous Improvement
- Develop evaluation frameworks and benchmarking systems that measure model quality, throughput, and cost.
- Use evaluation data to continuously improve production AI systems.
Cross-Functional Engineering
- Partner with infrastructure engineers to improve GPU scheduling and fleet utilization.
- Collaborate with product teams to prioritize and deliver new AI capabilities.
Qualifications
- 3+ years of experience designing, deploying, and operating production AI or machine learning systems.
- Hands-on experience serving large language models using vLLM, SGLang, TensorRT-LLM, or comparable inference platforms.
- Experience with reinforcement learning or post-training techniques, including GRPO, PPO, DPO, SFT, or equivalent approaches.
- Strong Python programming skills.
- Strong PyTorch experience.
- Working knowledge of GPU execution, including batching, memory optimization, and CUDA fundamentals.
- Ability to work independently with minimal supervision while operating effectively in ambiguous, fast-moving environments.
- Experience with slime, prime-rl, the Verifiers library, or Megatron-LM is preferred.
- Experience with distributed training technologies, including FSDP and tensor, pipeline, or data parallelism.
- Experience with quantization, speculative decoding, or production inference optimization.
- Experience building verifiable inference systems or large-scale distributed platforms.
- Experience with Kubernetes and containerized deployments.
- Familiarity with GPU orchestration platforms such as Ray, SkyPilot, or Slurm.
- A public GitHub profile is required.
- Your GitHub profile must demonstrate at least one year of development activity.
- Applications without a qualifying public GitHub profile will not be considered.
Benefits
- Health, dental, and vision insurance.
- Flexible paid time off.
- Professional development budget.
- Conference travel budget.
- Remote-first work environment.
- Regular company off-site events.
- High-trust, high-performance engineering culture.
Looking for more opportunities?
View All Jobs