Member of Technical Staff — Training
Quick facts
- Pay
- $200,000 to $400,000, plus equity
- Type
- Full-time
- Location
- Palo Alto, CA
- Experience
- 3+ years
- Industry
- Information Technology — 671 open jobs
- Visa in listing
- H-1B
- Apply
- Free to apply — never pay anyone for a job offer or visa sponsorship
Job description
About the Role
As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training — spanning training, inference, and orchestration, with a focus on the performance, correctness, scalability, and reliability of workloads running across large GPU clusters.
This role suits engineers who move fluidly across modeling recipes, complex infrastructure, and low-level systems, identify bottlenecks in distributed workloads, and translate experimental requirements into robust software.
In This Role, You Will
- Design, build, and operate distributed training, rollout, and orchestration systems for large-scale LLM and multimodal post-training across multi-GPU, multi-node environments.
- Profile and optimize performance across the full-stack — model implementation, parallelism strategies, communication libraries, and GPU kernels — to improve throughput, latency, memory efficiency, hardware utilization, and cost.
- Investigate numerical correctness and low-precision issues in distributed training and inference, including train–inference consistency for reinforcement learning.
- Improve the reliability of long-running workloads through checkpointing, fault recovery, observability, and operational tooling.
- Build supporting infrastructure for reinforcement learning and agentic post-training, including asynchronous rollout, trajectory collection, sandboxed execution, evaluation harnesses, and data pipelines.
- Contribute to open-source training and inference systems, including Miles and SGLang, and partner with researchers to turn experimental requirements into production systems.
Minimum Qualifications
- 3+ years of experience building or operating distributed machine learning systems, large-scale training infrastructure, or high-performance inference systems.
- Hands-on experience with post-training systems, training backends, or inference systems for large language models (e.g., Megatron-LM, FSDP, SGLang, TensorRT-LLM, vLLM).
- Experience in at least two of the following areas:
- Performance, efficiency, and scalability of multi-GPU, multi-node workloads
- Numerical correctness or low precision
- Stability, reliability, or fault tolerance
- Post-training algorithm recipes and orchestration infrastructure for large training runs
- Multimodal training or inference, including vision-language models and multimodal generation
- Agent infrastructure, including sandboxes, harnesses, and eval systems
- Building and maintaining open-source projects widely adopted in industry and academia
Preferred Qualifications
- Familiarity with RL algorithms such as PPO, GRPO, and their variants, and experience applying them in large-scale post-training.
- Experience with modern post-training frameworks (e.g., Miles, slime, AReaL, verl, Prime-RL).
- Key open-source contributions to training or inference frameworks (e.g., SGLang, vLLM, Megatron-LM).
- GPU kernel development (e.g., CUDA, Triton, CUTLASS) or communication-layer optimization (e.g., NCCL, RDMA, NVLink/NVSwitch).
- Experience training or serving models at very large scale (e.g., Mixture-of-Experts models on clusters of thousands of GPUs).
- Top-tier publications in ML systems or other systems fields.
Even if you don't meet every qualification above, we encourage you to apply — we care most about demonstrated ability to build and reason about large-scale systems.
About RadixArk
RadixArk builds open-source and production infrastructure for large language models and multimodal post-training. Our systems — including Miles, an enterprise-grade reinforcement learning training framework, and SGLang, a widely deployed high-performance LLM inference engine — power distributed post-training across clusters of 10k–100k+ GPUs.
Compensation
Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 to $400,000, plus equity.
Benefits include a 401(k) plan and unlimited PTO.
RadixArk sponsors employment visas (e.g., H-1B, O-1) for eligible candidates.
Equal Opportunity
RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Visa sponsorship record: RadixArk, Inc.
Sponsors actively| Fiscal year | H-1B LCAs | USCIS approvals | USCIS denials | PERM |
|---|---|---|---|---|
| FY2026 | 18 | 8 | 0 | — |
✔ RadixArk, Inc. filed 15 H-1B applications for this kind of role (Software Developers) in the period above.
Roles they sponsor most
- member of technical15 LCAs · median $188k
- member of technical1 LCA · median $209k
- member of technical1 LCA · median $174k
- product engineer1 LCA · median $180k
Which visas can work for this job
Occupation: Software Developers (SOC 15-1252).
- Mentioned in the listing
H-1B
- H-1B cap-exempt employer
Regular cap-subject employer — a new H-1B needs to win the lottery in March (unless you already hold cap-counted H-1B status).
- TN (citizens of Canada and Mexico)
This kind of role may fit the USMCA profession “Computer Systems Analyst (or Engineer, for engineering-degree holders)” — it depends on the actual duties. No lottery, no cap; you need the matching degree or license.
- E-3 (Australia) and H-1B1 (Chile, Singapore)
Same degree requirement as H-1B, but no lottery and a separate quota that is rarely filled.
- O-1 (extraordinary ability)
For candidates with awards, publications, press, a high salary or critical roles at distinguished organizations. No cap, no lottery; the employer files a petition.
- STEM OPT extension (F-1 students)
Requires an E-Verify employer. We didn't find this company in the E-Verify list — ask HR.
Salary vs prevailing wage
Level II (qualified)Software Developers · San Jose-Sunnyvale-Santa Clara, CA. Annual prevailing wages set by the Department of Labor (OFLC).
| Level | Prevailing wage | H-1B lottery odds* |
|---|---|---|
| Level I | $152,797 | ~15% |
| Level II | $187,075 | ~31% |
| Level III | $221,374 | ~46% |
| Level IV | $255,653 | ~61% |
This job pays $200,000–$400,000 a year — that's Level II (qualified). In the wage-weighted H-1B lottery a Level II registration gets 2 entries; estimated selection chance about 31%.
* Odds are DHS projections for the FY2027 wage-weighted lottery (actual results vary by year and employer). Since the FY2027 cap season the H-1B lottery is weighted by wage level: Level I = 1 entry, II = 2, III = 3, IV = 4. The level is set by the offered wage against the prevailing wage for the occupation and worksite. Separately, a $100,000 fee for new H-1B petitions for workers outside the US was announced in 2025; as of September 2026 a federal court ruling keeps it unenforceable while appeals continue — check the current status.
Sources: U.S. Department of Labor OFLC disclosure data (H-1B/H-1B1/E-3 LCA, PERM), OFLC prevailing wage data, USCIS H-1B Employer Data Hub, E-Verify participating employers. Data loaded: LCA FY2024–FY2026, PERM, USCIS Data Hub, OFLC wages; updated 2026-10-03. Employers are matched by name, so records of companies with similar names can occasionally be mixed up. This is general information, not legal advice — talk to an immigration attorney about your case.