Eligible for Health, Dental, Vision Not eligible for Visa sponsorship
Job Description:
Build, deploy, maintain, and optimize high-performance OpenSearch/Elasticsearch environments on Kubernetes supporting mission-critical cloud services.
Day-to-Day Responsibilities:
Architect, build, deploy, and maintain OpenSearch/Elasticsearch clusters on Kubernetes Monitor cluster health, node performance, indexing throughput, search latency, shard allocation, replication, and storage Troubleshoot production issues across infrastructure, platform, and application layers Lead incident response, root cause analysis, and post-incident remediation Handle platform installations, upgrades, patching, backup/restore, and disaster recovery Automate testing, deployment, scaling, recovery, and operational workflows Build and maintain CI/CD pipelines Support capacity planning across compute, memory, storage, and network resources Manage log ingestion, index management, retention policies, and search performance Develop monitoring, alerting, and operational runbooks Participate in an on-call rotation and occasional weekend/after-hours releases
Minimum Requirements:
8 years of SRE, DevOps, Platform Engineering, or related experience Expert-level Kubernetes experience in complex production environments Hands-on experience building, deploying, and maintaining OpenSearch or Elasticsearch clusters on Kubernetes in production Strong OpenSearch/Elasticsearch cluster administration, architecture, performance tuning, scaling, upgrades, and troubleshooting Experience with index design, shard/replica strategy, cluster sizing, snapshot/restore, and disaster recovery Strong Linux experience Experience with Concourse CI/CD pipelines Kafka and ZooKeeper experience Strong Git and automation/scripting experience Experience supporting distributed, highly available production systems Hands-on production incident response, RCA, and on-call experience Experience working within highly secure, complex enterprise environments
Preferred Qualifications:
AWS experience (EC2, S3, Route 53, CloudWatch, IAM, VPC, RDS) Cloud Foundry / PCF experience Terraform, Jenkins, and/or Chef Prometheus and Grafana Log ingestion and index lifecycle/retention management Capacity forecasting, performance benchmarking, and resilience testing SaaS / multi-tenant security experience