Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
M
Meta
Software Engineer, SystemML - AI Networking
Career Insights for Network Engineer / Architect
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on California data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Network Engineer or Architect designs and builds computer network systems, including software and hardware. Runs program and system tests, solves technical problems and maintains the network system. Designs and analyzes computer network models.
$146,860 / year median in California
-16% projected decline
Job Description
Software Engineer, SystemML
At the high level, the team aims to enable Meta-wide ML products and innovations to leverage our large-scale GPU training and inference fleet through an observable, reliable and high-performance distributed AI/GPU communication stack. Currently, one of the team's focus is on building customized features, software benchmarks, performance tuners and software stacks around NCCL and PyTorch to improve the full-stack distributed ML reliability and performance (e.g. Large-Scale GenAI/LLM training) from the trainer down to the inter-GPU and network communication layer. And we are seeking engineers to work on the space of GenAI/LLM scaling reliability and performance. Software Engineer, SystemML•
- AI Networking Meta
- 4.0 Menlo Park, CA Job Details $154,003
- $217,000 a year 4 hours ago Qualifications C Leading team collaboration initiatives Cross-functional team management Cross-functional communication Full Job Description In this role, you will be a member of the AI Networking Software team and part of the bigger DC networking organization.
At the high level, the team aims to enable Meta-wide ML products and innovations to leverage our large-scale GPU training and inference fleet through an observable, reliable and high-performance distributed AI/GPU communication stack. Currently, one of the team's focus is on building customized features, software benchmarks, performance tuners and software stacks around NCCL and PyTorch to improve the full-stack distributed ML reliability and performance (e.g. Large-Scale GenAI/LLM training) from the trainer down to the inter-GPU and network communication layer. And we are seeking engineers to work on the space of GenAI/LLM scaling reliability and performance. Software Engineer, SystemML•