Manager, Field Marketing - US Majors Financial Services

Bellevue, WA (Unknown)

$236K/yr – $310K/yrManager5+ years expFull time

Posted 8 days ago

Job Summary

Own and optimize high-performance LLM inference systems across distributed serving, runtime, and GPU kernels, driving latency, throughput, memory efficiency, and scalability while coordinating with researchers, infrastructure, and product teams.

  • Lead distributed serving, runtime, and kernel optimization efforts.
  • Bridge research innovations to production through collaboration with model, infra, and product teams.

Job Description

Responsibilities

  • Design and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels.
  • Develop novel techniques to improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.
  • Explore advanced inference techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, KV-cache management, quantization, and communication optimization.
  • Develop adaptive and intelligent inference systems that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments.
  • Apply AI-driven and AI-native approaches to systems engineering, including automated profiling, bottleneck identification, configuration search, code generation, experimentation, runtime strategy selection, debugging, and performance tuning.
  • Independently identify high-impact performance and systems problems, formulate hypotheses, prototype solutions, and drive promising ideas from research through production.
  • Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism.
  • Develop efficient approaches for multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling.
  • Analyze and optimize GPU kernels and operators for attention, MoE, communication, and other performance-critical model components.
  • Explore model-system co-design, including model or post-training techniques that unlock substantially more efficient inference.

Requirements

  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related field. A Master’s degree or PhD is preferred.
  • 5+ years of experience in one or more of the following areas: LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.
  • Strong understanding of modern LLM inference architectures and the performance tradeoffs involved in serving large-scale models.
  • Hands-on experience with modern LLM inference and serving frameworks, such as vLLM, SGLang, TensorRT-LLM, or similar systems.
  • Experience designing, extending, or optimizing inference runtimes, including areas such as scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving.
  • Strong understanding of GPU architectures and experience with CUDA, Triton, or similar GPU programming environments.
  • Experience with performance-oriented libraries and frameworks such as CUTLASS, cuBLAS, cuDNN, or related technologies.
  • Experience profiling and diagnosing end-to-end system performance using Nsight Systems, Nsight Compute, or equivalent tools.
  • Demonstrated ability to operate as an independent problem identifier and solver—recognizing important problems with limited direction, defining the right technical questions, and driving solutions through ambiguity.
  • Strong ability to work across model, runtime, distributed system, and hardware layers and reason about end-to-end performance tradeoffs.
Snowflake logo

Snowflake

3.9 Glassdoor

**Snowflake is proud to be the Official Data Collaboration Provider for LA28 and Team USA.**Snowflake delivers the AI Data Cloud — a global network where thousands of organizations mobilize data with near-unlimited scale, concurrency, and performance. Inside the AI Data Cloud, organizations unite their siloed data, easily discover and securely share governed data, and execute diverse analytic workloads. Wherever data or users live, Snowflake delivers a single and seamless experience across multiple public clouds. Snowflake’s platform is the engine that powers and provides access to the AI Data Cloud, creating a solution for data warehousing, data lakes, data engineering, data science, data application development, and data sharing. Join Snowflake customers, partners, and data providers already taking their businesses to new frontiers in the AI Data Cloud.

Founded in 2012
Cloud, Iowa, USA
11K employees (633 in marketing)
74% recommend to a friend
85% CEO approval

Salaries

Posted range

$236K/yr – $310K/yr

Taken straight from the job posting.

See all Snowflake salaries

Funding

Total Funding
$2.4B
Last Raise
$376.1M
Post IPO Equity, 1 months ago

Funding History

  • Feb 2026
    Post IPO Equity • $376.1M
  • Apr 2022
    Post IPO Equity • $621.5M
  • Feb 2020
    Series G • $479M

Traffic Signals

Monthly Visitors
4.6M
Traffic Source Mix
Search
27.9%
Direct
52.7%
Referral
16.3%
Social
1.9%
Paid
1.1%

Headcount Trend

Current headcount: ~10.7K

Recent News