Machine Learning Engineer
Vertex Agility
New York, United States Full Time Engineering Jobs United States New
Job Description
We are looking for an experienced ML Engineer to join an established Machine Learning Engineering team building and maintaining production AI systems. The role focuses on developing LLM-powered applications, scalable ML inference pipelines, and robust MLOps infrastructure. You'll work closely with Engineering and Data teams to deliver production-ready AI solutions using both hosted LLM APIs and open-source models.
Key Responsibilities
- Design, develop, and maintain production-grade LLM-powered applications, including chat interfaces and MCP servers
- Build and optimize AI systems using OpenAI, Anthropic Claude, and open-source LLMs
- Develop agentic AI workflows including function calling, tool use, and multi-step reasoning pipelines
- Deploy, optimize, and maintain ML classification models for scalable production environments
- Build reliable inference pipelines with efficient GPU/CPU utilization and performance optimization
- Develop and maintain reusable MLOps infrastructure across AI workloads
- Implement CI/CD pipelines, automated testing, and deployment processes for ML systems
- Monitor production AI systems through observability, logging, metrics, and alerting
- Write SQL queries to validate model outputs and monitor data quality
- Troubleshoot and resolve production incidents involving ML and LLM applications
- Collaborate with Data Scientists and Software Engineers on architecture, feature enhancements, and deployment improvements
- Participate in technical design discussions and contribute to improving engineering best practices.
Key Skills & Experience
- Strong hands-on experience with Python and production software engineering practices
- Proven experience building and deploying LLM-powered applications
- Experience with OpenAI and Anthropic Claude APIs in production environments
- Hands‑on experience implementing Agentic AI, Function Calling, Tool Use, or MCP (Model Context Protocol)
- Experience deploying and serving open‑source LLMs (Hugging Face or similar)
- Strong understanding of GPU acceleration, scalable inference, and ML performance optimization
- Experience with AWS services (Batch, ECS, Lambda, SQS, or similar)
- Strong SQL skills and experience working with data warehouses
- Experience building and maintaining CI/CD pipelines for ML workloads
- Knowledge of observability, logging, monitoring, and alerting for production ML systems
- Strong Git, debugging, documentation, and stakeholder communication skills.
Nice to Have
- Experience with Natural Language Processing (NLP)
- Experience with Vector Databases (Pinecone, Weaviate, pgvector)
- Hands‑on experience implementing Retrieval‑Augmented Generation (RAG) architectures
- Experience working across both hosted LLM APIs and open-source AI models.
Posted July 23, 2026