FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Staff Machine Learning Engineer – ML Frameworks
AdobeStaff ML Engineer building Kubernetes, Python, and GPU infrastructure for Adobe Firefly generative AI models. Improving distributed training, orchestration, scalability, and experimentation.
Posted 9/2/2026full-timeSan Jose • California, Washington • 🇺🇸 United StatesLead💰 $172,500 - $306,625 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and maintaining AI/ML infrastructure solutions, with a strong focus on Python, Kubernetes, and AWS. Proven ability to enhance distributed training frameworks and optimize resource utilization for large-scale AI model deployment.
Highest-signal resume keywords
Python ProficiencyKubernetes ExperienceDistributed PyTorch KnowledgeAI/ML Infrastructure DevelopmentGPU Resource Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AI/ML Infrastructure SolutionsDistributed Training FrameworksModel ServingOrchestration and SchedulingAutoML ToolsPython DevelopmentCritical ThinkingAnalytical Problem-SolvingQuantitative AnalysisMachine Learning
Soft Skills
Excellent CommunicationRelationship SkillsTeam Collaboration
Tools & Technologies
KubeFlowMLFlowRaySageMakerPyTorch DistributedMPIMegatronHorovod
Certifications & Qualifications
PhD in Computer ScienceMaster’s in Computer Science
Industry Keywords
AI ModelsMachine Learning ResearchInfrastructure PracticesResource UtilizationScalability
Tech Stack
Tools & technologiesAWSCloudKubernetesPythonPyTorchRay
About the role
Key responsibilities & impact- Design, develop, and maintain robust AI/ML infrastructure solutions supporting training and deployment of large-scale AI models using Kubernetes and Python on AWS cloud
- Implement and improve distributed training frameworks leveraging GPUs to improve performance and scalability
- Improve resiliency, elasticity, data loading, and out-of-the-box support for FSDP and model parallelism
- Improve orchestration and scheduling to train better models
- Scale the number of jobs and enable faster experimentation with AutoML and similar tools
- Collaborate with data scientists and ML researchers to streamline model training pipelines and ensure efficient resource utilization
- Drive innovation in infrastructure practices supporting machine learning research and development
Requirements
What you’ll need- PhD or Master’s in computer science or related field and 5+ years of hands-on industry experience
- Proven proficiency with Python and developing systems, frameworks and SDKs
- Experience with infrastructure and understanding of model serving, training, orchestration, and management of GPU resources
- Experience with machine learning and distributed PyTorch
- Strong critical thinking, analytical and quantitative problem-solving ability
- Excellent communication, relationship skills and a strong teammate
- Experience with KubeFlow, MLFlow, Ray, SageMaker, or similar (added plus)
- Experience with PyTorch distributed, MPI, Megatron, Horovod and other AI training frameworks (added plus)
Benefits
Comp & perks- Annual Incentive Plan (AIP) for short-term incentives
- Certain roles may be eligible for a new hire equity award
- Comprehensive benefits programs
- Equal Employment Opportunity protections
- Disability accommodations during the recruiting and application process