Member of Technical Staff, Inference Systems
Quick facts
- Pay
- $230K–$350K + equity
- Type
- Full-time
- Location
- Palo Alto, California
- Experience
- 2–10 years
- Skills
- Python, C++, Rust, PyTorch
- Industry
- Information Technology — 674 open jobs
- Visa in listing
- H-1B, OPT
- Apply
- Free to apply — never pay anyone for a job offer or visa sponsorship
Job description
About the Company
Our client is a seed-stage, stealth-mode startup building a high-performance AI inference platform from the ground up. The founding team brings deep AI-infrastructure experience and is tackling the hardest problems in LLM serving — scheduling, KV cache management, request routing, and the runtime systems that power model inference at scale — with Rust at the core of the stack. This is a ground-floor opportunity to shape the architecture of a fast, reliable inference system without legacy constraints.
Recently founded · Small founding team · Industry: AI infrastructure / LLM inference
The Role
You'll join a small, fast-moving team building a new inference system from scratch. This is a role for a systems engineer who lives and breathes inference internals — attention, KV cache, batching, scheduling — and wants to own the whole stack rather than a narrow slice.
What you'll be doing
Build a new inference runtime from scratch in Rust, owning batching, scheduling, request routing, and the full serving stack.
Design and implement KV cache management, prefix caching, and optimizations that cut latency and cost per token.
Scale serving across GPUs and nodes, tackling multi-GPU and multi-node challenges directly.
Profile, benchmark, and ship performance improvements across the entire inference pipeline.
Work closely with the founding team on the core architectural decisions that define the platform.
Tech stack: Rust, Python, PyTorch, C++, Go, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, NCCL
Requirements
2–10 years of experience as a backend or distributed systems engineer.
Hands-on experience building, operating, or optimizing LLM inference or serving systems at the engine, router, or runtime layer — beyond just calling hosted APIs.
A deep working knowledge of transformer inference internals: attention, KV cache, batching, scheduling, and where the real bottlenecks are.
Performance-critical backend or distributed systems experience where latency, throughput, and cost were first-order concerns.
Hands-on time with a production inference engine (vLLM, SGLang, or TensorRT-LLM) plus strong systems-language skills (Rust, C++, Go, or systems-level Python/PyTorch).
If you haven't used Rust yet, you should be able to become productive in it within a few weeks of joining.
Nice to Haves
Time on an inference team at a model provider, accelerator vendor, or research lab, or open-source contributions to vLLM, SGLang, or Dynamo. A CS or systems degree from a strong program. Production Rust experience, CUDA/Triton kernel work, multi-GPU or multi-node serving (NCCL, NVLink, RDMA), prefix caching, speculative decoding, or prefill/decode disaggregation.
Why Join
Own the architecture of a brand-new inference system with no legacy constraints.
Solve the hardest problems in LLM serving, from KV cache management to request routing.
A performance engineer's dream: latency and cost per token are the whole game.
High visibility on a small, elite team where your code powers the core engine.
Details
Location: Palo Alto, CA
Work policy: Full-time, on-site five days a week
Compensation: $230K–$350K + equity
Visa sponsorship: Open to visa transfers (OPT, H-1B) and new sponsorships (new H-1B, TN)
Employment type: Full-time
Visa sponsorship record
No public sponsorship records foundWe didn't find H-1B, H-1B1, E-3 or green card (PERM) filings under the name David Joseph & Company in the Department of Labor and USCIS data. That doesn't mean the job can't be sponsored — the company may file under a different legal name, be new to sponsorship, or sponsor a visa that isn't in these datasets (for example H-2B or J-1).
Tip: ask the recruiter early, in writing, which visa they sponsor and whether they cover the legal and filing fees.
Which visas can work for this job
Occupation: Software Developers (SOC 15-1252).
- Mentioned in the listing
H-1B, OPT
- H-1B cap-exempt employer
Regular cap-subject employer — a new H-1B needs to win the lottery in March (unless you already hold cap-counted H-1B status).
- TN (citizens of Canada and Mexico)
This kind of role may fit the USMCA profession “Computer Systems Analyst (or Engineer, for engineering-degree holders)” — it depends on the actual duties. No lottery, no cap; you need the matching degree or license.
- E-3 (Australia) and H-1B1 (Chile, Singapore)
Same degree requirement as H-1B, but no lottery and a separate quota that is rarely filled.
- O-1 (extraordinary ability)
For candidates with awards, publications, press, a high salary or critical roles at distinguished organizations. No cap, no lottery; the employer files a petition.
- STEM OPT extension (F-1 students)
Requires an E-Verify employer. We didn't find this company in the E-Verify list — ask HR.
Salary vs prevailing wage
Level III (experienced)Software Developers · San Jose-Sunnyvale-Santa Clara, CA. Annual prevailing wages set by the Department of Labor (OFLC).
| Level | Prevailing wage | H-1B lottery odds* |
|---|---|---|
| Level I | $152,797 | ~15% |
| Level II | $187,075 | ~31% |
| Level III | $221,374 | ~46% |
| Level IV | $255,653 | ~61% |
This job pays $230,000–$350,000 a year — that's Level III (experienced). In the wage-weighted H-1B lottery a Level III registration gets 3 entries; estimated selection chance about 46% — better than average.
* Odds are DHS projections for the FY2027 wage-weighted lottery (actual results vary by year and employer). Since the FY2027 cap season the H-1B lottery is weighted by wage level: Level I = 1 entry, II = 2, III = 3, IV = 4. The level is set by the offered wage against the prevailing wage for the occupation and worksite. Separately, a $100,000 fee for new H-1B petitions for workers outside the US was announced in 2025; as of September 2026 a federal court ruling keeps it unenforceable while appeals continue — check the current status.
Sources: U.S. Department of Labor OFLC disclosure data (H-1B/H-1B1/E-3 LCA, PERM), OFLC prevailing wage data, USCIS H-1B Employer Data Hub, E-Verify participating employers. Data loaded: LCA FY2024–FY2026, PERM, USCIS Data Hub, OFLC wages; updated 2026-10-03. Employers are matched by name, so records of companies with similar names can occasionally be mixed up. This is general information, not legal advice — talk to an immigration attorney about your case.