Member of Technical Staff - Compilers at Gimlet Media | Job-Scouts.com

Member of Technical Staff - Compilers

Gimlet Media
full-time lead San Francisco, CA
This position is sourced from Gimlet Media's career page . Apply through Job-Scouts to track your application status.

Job Description

About us

Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.

We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.

We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.

About the role

As a Member of Technical Staff, you will build the compiler infrastructure that determines how AI workloads are represented, optimized, and executed across hardware with different architectures, performance characteristics, and memory systems.

This is compiler engineering at the boundary of ML systems and distributed execution. The problems do not end when code is generated: compiler decisions interact directly with scheduling, communication, memory movement, kernel execution, and serving performance.

You will work across intermediate representations, graph transformations, lowering, execution planning, and runtime interfaces. You will develop strategies for partitioning computation across devices, bring new models and accelerator architectures onto the platform, and partner with ML systems, kernel, and distributed systems engineers to improve end-to-end execution.

Our work on Corsair and low-latency speculative decoding is one example of the problems this team tackles.

What success looks like

In your first 12-18 months, you will:

• Improve the latency, throughput, and efficiency of production inference workloads

• Design execution strategies for partitioning workloads across heterogeneous hardware

• Develop compiler optimizations spanning IR transformations, scheduling, memory movement, and kernel orchestration

• Enable new models, accelerator architectures, and serving techniques to run efficiently in production

You may be a good fit if you have

• Experience building compiler, runtime, or execution infrastructure

• Experience with IR transformations, compiler passes, lowering, or code generation

• Strong systems and performance-engineering fundamentals

• The ability to reason about execution behavior, memory systems, scheduling, and hardware efficiency

• Strong C++ and/or Python skills

• A bachelor’s degree in a relevant field or equivalent practical experience

Strong candidates may also have

• Experience with MLIR, LLVM, XLA, TVM, Triton, or similar compiler/runtime infrastructure

• Experience optimizing ML inference or serving workloads

• Familiarity with runtime systems, kernel dispatch, launch APIs, or memory allocators

• Experience working with GPUs, AI accelerators, or heterogeneous hardware systems

• Experience profiling and debugging performance-critical systems

• Familiarity with scheduling, partitioning, or kernel-level optimizations

Why join now?

Gimlet is expanding from its core technology into a production neocloud spanning new hardware, customers, and data centers.

• Solve hard problems.

• Own meaningful work.

• Build for production.

• Help define what’s next.

Agency Policy: Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.

Requirements

Department: Research and Development
Team: Engineering

Location

San Francisco, CA
View on Google Maps

Similar open positions