LLM Inference & GPU Systems Engineer - Charlotte, NC (Onsite)

0 views
Skip to first unread message

Sid K

unread,
10:01 AM (1 hour ago) 10:01 AM
to

Hi Folks,


My client is looking for LLM Inference & GPU Systems Engineer for 12 Month Contract role based in Charlotte, NC (Onsite)

 

Job Title: LLM Inference & GPU Systems Engineer

Location: Charlotte, NC (Onsite)
Duration: 12 Month Contract
Rate: $55/hr. on C2C Max


Role Overview

We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.

 

Key Responsibilities

NVIDIA GPU Runtime Optimization: Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.

Inference Serving: Deploy and manage inference engines including vLLM and TensorRT-LLM.

Hardware Utilization: Optimize GPU throughput tuning, batching strategies, and latency optimization. Manage workload orchestration using RunAI and Kubernetes GPU orchestration.Model Lifecycle Management: Oversee the complete Hugging Face model lifecycle, including model onboarding, deployment, and retirement.

Platform Operations: Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.

 

Required Qualifications

1-3 years experience as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.

Built a multi-agent AI analytics platform using DeepAnalyze-8B, Gemini and vLLM with LLM inference and execution workflows.

Experience developing scalable ML pipelines and deployment-ready AI architectures at Dassault Systèmes.

Strong foundation in REST APIs, backend development, ML pipelines, and containerized AI applications.

Thanks
Sid

Reply all
Reply to author
Forward
0 new messages