Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Adobe

Staff Machine Learning Engineer – ML Frameworks

Adobe

Staff ML Engineer building Kubernetes, Python, and GPU infrastructure for Adobe Firefly generative AI models. Improving distributed training, orchestration, scalability, and experimentation.

Posted 9/2/2026full-timeSan Jose • California, Washington • 🇺🇸 United StatesLead💰 $172,500 - $306,625 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and maintaining AI/ML infrastructure solutions, with a strong focus on Python, Kubernetes, and AWS. Proven ability to enhance distributed training frameworks and optimize resource utilization for large-scale AI model deployment.

Highest-signal resume keywords
Python ProficiencyKubernetes ExperienceDistributed PyTorch KnowledgeAI/ML Infrastructure DevelopmentGPU Resource Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
AI/ML Infrastructure SolutionsDistributed Training FrameworksModel ServingOrchestration and SchedulingAutoML ToolsPython DevelopmentCritical ThinkingAnalytical Problem-SolvingQuantitative AnalysisMachine Learning
Soft Skills
Excellent CommunicationRelationship SkillsTeam Collaboration
Tools & Technologies
KubeFlowMLFlowRaySageMakerPyTorch DistributedMPIMegatronHorovod
Certifications & Qualifications
PhD in Computer ScienceMaster’s in Computer Science
Industry Keywords
AI ModelsMachine Learning ResearchInfrastructure PracticesResource UtilizationScalability

Tech Stack

Tools & technologies
AWSCloudKubernetesPythonPyTorchRay

About the role

Key responsibilities & impact
  • Design, develop, and maintain robust AI/ML infrastructure solutions supporting training and deployment of large-scale AI models using Kubernetes and Python on AWS cloud
  • Implement and improve distributed training frameworks leveraging GPUs to improve performance and scalability
  • Improve resiliency, elasticity, data loading, and out-of-the-box support for FSDP and model parallelism
  • Improve orchestration and scheduling to train better models
  • Scale the number of jobs and enable faster experimentation with AutoML and similar tools
  • Collaborate with data scientists and ML researchers to streamline model training pipelines and ensure efficient resource utilization
  • Drive innovation in infrastructure practices supporting machine learning research and development

Requirements

What you’ll need
  • PhD or Master’s in computer science or related field and 5+ years of hands-on industry experience
  • Proven proficiency with Python and developing systems, frameworks and SDKs
  • Experience with infrastructure and understanding of model serving, training, orchestration, and management of GPU resources
  • Experience with machine learning and distributed PyTorch
  • Strong critical thinking, analytical and quantitative problem-solving ability
  • Excellent communication, relationship skills and a strong teammate
  • Experience with KubeFlow, MLFlow, Ray, SageMaker, or similar (added plus)
  • Experience with PyTorch distributed, MPI, Megatron, Horovod and other AI training frameworks (added plus)

Benefits

Comp & perks
  • Annual Incentive Plan (AIP) for short-term incentives
  • Certain roles may be eligible for a new hire equity award
  • Comprehensive benefits programs
  • Equal Employment Opportunity protections
  • Disability accommodations during the recruiting and application process