Member of Technical Staff, Inference Systems
Quick facts
- Pay
- $230,000 - $350,000 a year
- Type
- Full-time
- Location
- Palo Alto, CA
- Experience
- 2–10 years
- Skills
- Python, C++, Rust, PyTorch
- Industry
- Information Technology — 671 open jobs
- Visa in listing
- H-1B, OPT
- Apply
- Free to apply — never pay anyone for a job offer or visa sponsorship
Job description
About the Company
Our client is a seed-stage, stealth-mode startup building a high-performance AI inference platform from the ground up. The founding team brings deep AI-infrastructure experience and is tackling the hardest problems in LLM serving — scheduling, KV cache management, request routing, and the runtime systems that power model inference at scale — with Rust at the core of the stack. This is a ground-floor opportunity to shape the architecture of a fast, reliable inference system without legacy constraints.
- Recently founded
- Small founding team
- Industry: AI infrastructure / LLM inference
The Role
You'll join a small, fast-moving team building a new inference system from scratch. This is a role for a systems engineer who lives and breathes inference internals — attention, KV cache, batching, scheduling — and wants to own the whole stack rather than a narrow slice.
What you'll be doing
- Build a new inference runtime from scratch in Rust, owning batching, scheduling, request routing, and the full serving stack.
- Design and implement KV cache management, prefix caching, and optimizations that cut latency and cost per token.
- Scale serving across GPUs and nodes, tackling multi-GPU and multi-node challenges directly.
- Profile, benchmark, and ship performance improvements across the entire inference pipeline.
- Work closely with the founding team on the core architectural decisions that define the platform.
Tech stack: Rust, Python, PyTorch, C++, Go, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, NCCL
Requirements
- 2–10 years of experience as a backend or distributed systems engineer.
- Hands-on experience building, operating, or optimizing LLM inference or serving systems at the engine, router, or runtime layer — beyond just calling hosted APIs.
- A deep working knowledge of transformer inference internals: attention, KV cache, batching, scheduling, and where the real bottlenecks are.
- Performance-critical backend or distributed systems experience where latency, throughput, and cost were first-order concerns.
- Hands-on time with a production inference engine (vLLM, SGLang, or TensorRT-LLM) plus strong systems-language skills (Rust, C++, Go, or systems-level Python/PyTorch).
- If you haven't used Rust yet, you should be able to become productive in it within a few weeks of joining.
Nice to Haves
- Time on an inference team at a model provider, accelerator vendor, or research lab, or open-source contributions to vLLM, SGLang, or Dynamo.
- A CS or systems degree from a strong program.
- Production Rust experience, CUDA/Triton kernel work, multi-GPU or multi-node serving (NCCL, NVLink, RDMA), prefix caching, speculative decoding, or prefill/decode disaggregation.
Why Join
- Own the architecture of a brand-new inference system with no legacy constraints.
- Solve the hardest problems in LLM serving, from KV cache management to request routing.
- A performance engineer's dream: latency and cost per token are the whole game.
- High visibility on a small, elite team where your code powers the core engine.
Details
Location: Palo Alto, CA
Work policy: Full-time, on-site five days a week
Compensation: $230K–$350K + equity
Visa sponsorship: Open to visa transfers (OPT, H-1B) and new sponsorships (new H-1B, TN)
Employment type: Full-time
Visa sponsorship record
No public sponsorship records foundWe didn't find H-1B, H-1B1, E-3 or green card (PERM) filings under the name David Joseph & Company in the Department of Labor and USCIS data. That doesn't mean the job can't be sponsored — the company may file under a different legal name, be new to sponsorship, or sponsor a visa that isn't in these datasets (for example H-2B or J-1).
Tip: ask the recruiter early, in writing, which visa they sponsor and whether they cover the legal and filing fees.
Which visas can work for this job
Occupation: Software Developers (SOC 15-1252).
- Mentioned in the listing
H-1B, OPT
- H-1B cap-exempt employer
Regular cap-subject employer — a new H-1B needs to win the lottery in March (unless you already hold cap-counted H-1B status).
- TN (citizens of Canada and Mexico)
This kind of role may fit the USMCA profession “Computer Systems Analyst (or Engineer, for engineering-degree holders)” — it depends on the actual duties. No lottery, no cap; you need the matching degree or license.
- E-3 (Australia) and H-1B1 (Chile, Singapore)
Same degree requirement as H-1B, but no lottery and a separate quota that is rarely filled.
- O-1 (extraordinary ability)
For candidates with awards, publications, press, a high salary or critical roles at distinguished organizations. No cap, no lottery; the employer files a petition.
- STEM OPT extension (F-1 students)
Requires an E-Verify employer. We didn't find this company in the E-Verify list — ask HR.
Salary vs prevailing wage
Level III (experienced)Software Developers · San Jose-Sunnyvale-Santa Clara, CA. Annual prevailing wages set by the Department of Labor (OFLC).
| Level | Prevailing wage | H-1B lottery odds* |
|---|---|---|
| Level I | $152,797 | ~15% |
| Level II | $187,075 | ~31% |
| Level III | $221,374 | ~46% |
| Level IV | $255,653 | ~61% |
This job pays $230,000–$350,000 a year — that's Level III (experienced). In the wage-weighted H-1B lottery a Level III registration gets 3 entries; estimated selection chance about 46% — better than average.
* Odds are DHS projections for the FY2027 wage-weighted lottery (actual results vary by year and employer). Since the FY2027 cap season the H-1B lottery is weighted by wage level: Level I = 1 entry, II = 2, III = 3, IV = 4. The level is set by the offered wage against the prevailing wage for the occupation and worksite. Separately, a $100,000 fee for new H-1B petitions for workers outside the US was announced in 2025; as of September 2026 a federal court ruling keeps it unenforceable while appeals continue — check the current status.
Sources: U.S. Department of Labor OFLC disclosure data (H-1B/H-1B1/E-3 LCA, PERM), OFLC prevailing wage data, USCIS H-1B Employer Data Hub, E-Verify participating employers. Data loaded: LCA FY2024–FY2026, PERM, USCIS Data Hub, OFLC wages; updated 2026-10-03. Employers are matched by name, so records of companies with similar names can occasionally be mixed up. This is general information, not legal advice — talk to an immigration attorney about your case.