Job Title: Data Engineer – AWS OpenSearch
Location: Fort Mill, SC – Hybrid
Visa- GC and USC
Job Summary
We are seeking an experienced Data Engineer with strong expertise in AWS OpenSearch to design, develop, and maintain scalable data pipelines and search solutions. The ideal candidate will have hands-on experience with AWS data services, OpenSearch, Python, SQL, and cloud-native architectures to support analytics and real-time search capabilities.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using AWS services.
- Build and optimize AWS OpenSearch clusters for indexing, search, and analytics.
- Develop ETL/ELT pipelines using Python, SQL, and AWS services such as Glue, Lambda, and EMR.
- Design and maintain OpenSearch indexes, mappings, analyzers, and query performance.
- Integrate OpenSearch with upstream and downstream applications for real-time search and analytics.
- Work with structured and semi-structured data from multiple data sources.
- Optimize OpenSearch cluster performance, shard allocation, indexing, and query execution.
- Implement Infrastructure as Code (Terraform or CloudFormation) for AWS resources.
- Build monitoring and alerting using CloudWatch, Grafana, or Prometheus.
- Collaborate with Data Scientists, Data Analysts, and application teams to support data and search requirements.
- Implement security best practices, IAM policies, encryption, and access controls.
- Troubleshoot production issues and optimize system performance.
Required Skills
- 5+ years of experience as a Data Engineer.
- Strong experience with AWS OpenSearch (formerly Amazon Elasticsearch Service).
- Strong SQL and Python programming skills.
- Experience with AWS services:
- OpenSearch
- S3
- Glue
- Lambda
- IAM
- CloudWatch
- EC2
- EMR
- Step Functions
- Kinesis (preferred)
- Experience building ETL/ELT pipelines.
- Experience with Docker and Kubernetes (preferred).
- Knowledge of distributed search architecture, indexing strategies, and query optimization.
- Experience with Git, CI/CD, and Infrastructure as Code (Terraform or CloudFormation).
- Familiarity with Agile/Scrum methodologies.
Preferred Qualifications
- Experience with Apache Spark or PySpark.
- Experience with Kafka or Amazon MSK.
- Experience with Snowflake, Redshift, or Databricks.
- Experience working with large-scale distributed systems.
- AWS Certification (Solutions Architect, Data Analytics, or Developer) is a plus.
- Experience with observability tools such as Grafana, Prometheus, or Splunk.
Nice to Have
- Vector search and semantic search using OpenSearch.
- OpenSearch Dashboards/Kibana.
- Machine Learning integrations with OpenSearch.
- Experience with GenAI/RAG search architectures.
- Healthcare, Financial Services, or Retail domain experience.
Primary Technologies: AWS OpenSearch, Python, SQL, AWS Glue, Lambda, S3, Terraform, Docker, Kubernetes, CloudWatch, Git, CI/CD, Spark/PySpark.