We are looking for a Senior Reliability Engineer to strengthen the stability, scalability, and performance of technology systems that support construction and field operations. This position plays a key role in building dependable infrastructure across remote job sites, field offices, and cloud environments so teams can work efficiently with minimal disruption. The ideal candidate brings a strong SRE mindset, combines technical depth with practical problem-solving, and partners effectively with operational teams to maintain business continuity and system resilience.
Responsibilities:
- Build and enhance highly available infrastructure that supports office locations, remote field environments, networking needs, cloud services, and edge-based systems.
- Direct incident response efforts for service disruptions, coordinate restoration activities, and lead root cause investigations to prevent repeat issues.
- Create and maintain monitoring, alerting, and observability capabilities that improve visibility into system health, uptime, and application performance.
- Work closely with construction, engineering, and field personnel to ensure technology reliability aligns with project schedules, operational demands, and safety expectations.
- Implement automated infrastructure deployment and recovery processes using Infrastructure as Code and configuration management tools such as Terraform and Ansible.
- Establish service reliability targets, manage service level objectives, and use error budgets to guide operational decisions and continuous improvement.
- Strengthen the security posture of remote and field-deployed systems by improving hardening practices and secure access methods.
- Provide guidance to less experienced engineers and help foster a culture centered on reliability, accountability, and operational excellence.
- Identify process improvements that reduce inefficiencies, simplify support efforts, and improve overall service delivery.