Job opening
Senior LLMOps / MLOps Engineer
Filed under IT Services and IT Consulting
Full job description
GyanSys is looking for Senior LLMOps / MLOps Engineer to join one of our direct clients in Santa Clara, CA
Please see the details below and let me know if you are interested,
- 5-7 years of experience in MLOps, LLMOps, AI/ML Platform Engineering.
- Experience deploying, scaling, and monitoring production-grade GenAI/LLM applications.
- Hands-on experience with LLM Fine-Tuning using PEFT, SFT, CPT, LoRA, and QLoRA techniques.
- Experience with Azure AI Foundry, Azure OpenAI, Hugging Face, DeepSpeed, and PEFT.
- Knowledge of distributed training and multi-GPU environments.
- Experience with Agentic AI frameworks such as LangGraph, AutoGen, or CrewAI.
- Experience working with open-source LLMs such as Llama, Mistral, Gemma, or Qwen.
- Strong expertise in LLM Inferencing and Model Hosting using vLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving.
- Experience with Kubernetes, Docker, Azure ML, Databricks, and MLflow.
- Good understanding of RAG, Vector Databases, GPU Optimization, Quantization, KV Cache, PagedAttention, and Continuous/Dynamic Batching.
- Demonstrated hands-on experience building, deploying, troubleshooting, and optimizing production-grade LLM and GenAI solutions.
- Location: Santa Clara, CA (Onsite Mon-Fri)