Lead end-to-end solution
design across data ingestion, model training, inference, deployment, and
monitoring.
Architect LLM
applications (agents, summarization, classification) with RAG, evaluation
frameworks, safety controls, and guardrails.
Own MLOps/LLMOps
practices, including CI/CD for models, model registry, feature stores,
lineage tracking, observability, drift detection, and cost monitoring.
Choose the right cloud
and runtime strategy (managed services vs. self-hosted, GPU vs. CPU,
serverless vs. containerized).
Establish AI governance
standards, including PII handling, encryption, auditability, and
Responsible AI practices.
Collaborate with product
and business stakeholders to translate requirements into architectural
decisions and delivery plans.
Perform technical spikes
and POCs, benchmark models and infrastructure, and lead Architecture
Reviews.
Create and maintain
standards, patterns, and reusable components; mentor engineers across
teams.
Drive performance and
cost optimization initiatives, including throughput, latency, SLA/SLO
management, caching, quantization/distillation, and autoscaling.
Support vendor and
product evaluations, including cloud AI services, vector databases,
orchestration frameworks, and monitoring platforms.
Required Qualifications
Bachelor’s or master’s
degree in computer science, Engineering, Data Science, AI, or a related
field.
15+ years of overall
engineering experience, with at least 4+ years in AI/ML solution
architecture.
Proven experience
designing and deploying AI systems in production at scale (LLM and/or
classical ML).
Strong hands-on
proficiency in Python and at least one major cloud platform (AWS, Azure,
or GCP).
Must-Have Technical Skills
AI/ML & LLM Architecture
Designing LLM/RAG
systems, including retrieval pipelines, chunking strategies, embeddings,
reranking, prompt orchestration, response orchestration, evaluation, and
safety.
Deep understanding of
the model lifecycle, including fine-tuning, PEFT/LoRA, quantization,
distillation, latency optimization, and cost optimization.
Strong ML/NLP expertise,
including feature engineering, model selection, training,
cross-validation, experimentation, and testing.
MLOps / LLMOps
CI/CD for ML, including
model versioning, model promotion, feature stores, model registry, lineage
tracking, and drift detection.
Inference stacks
including PyTorch, TensorFlow, vLLM, TGI, ONNX, GPU orchestration,
autoscaling, and APM.
Pipelines and
orchestration frameworks such as Airflow, Kubeflow, and MLflow.