Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Insight Global

Senior Software Engineer, Reliability Engineering - REMOTE

Career Insights for Reliability Engineer

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on national data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

A Reliability Engineer evaluates and analyzes the reliability of manufacturing and production equipment and processes. Identifies pro-active maintenance requirements and process improvements. Works to maximize reliable performance and prevent equipment or process failures.

$119,450 / year median in the U.S.

+6% projected growth

Explore Career

Job Description

Job Description Senior Software Engineer, Reliability Engineering Position Summary Our Healthcare Client is seeking a Senior Software Engineer, Reliability Engineering who combines strong software development expertise with a passion for building highly reliable, scalable, observable, and operationally excellent systems. This role sits at the intersection of software engineering, cloud architecture, and site reliability engineering. You will design, build, and support customer-facing applications and APIs while ensuring reliability, scalability, security, performance, and operational excellence are built into every stage of the software lifecycle. Unlike traditional SRE or operations-focused roles, this position places equal emphasis on software engineering and reliability engineering. You will develop business capabilities, while helping teams improve observability, resiliency, deployment safety, and operational maturity. You will work closely with Product Engineering, Architecture, Platform Engineering, Security, and Product teams to enable rapid innovation without compromising system reliability. This role is designed to support mobility across engineering disciplines, allowing engineers to rotate between product development and reliability engineering teams. This is an engineering-first role with end-to-end production ownership, including participation in an on-call rotation for critical customer-facing applications and services. Hiring Philosophy We believe reliability is a shared engineering responsibility, not a separate operational function. The ideal candidate is a strong software engineer who enjoys building resilient systems, automating operational processes, improving developer productivity, and owning services throughout their entire lifecycle. Engineers in this role are expected to contribute to production ready code, influence system design, and help create a culture where reliability, scalability, and operational excellence are integral parts of software development. What You'll Do Software Engineering
  • Design, develop, test, deploy, and maintain cloud-native resilient distributed systems using APIs, microservices, messaging, caching, and event-driven architectures.
  • Build scalable and resilient services using Node.js and TypeScript
  • Participate in architecture and design discussions with a focus on scalability, resiliency, maintainability, security, and performance.
  • Apply modern software engineering practices, including domain-driven design, test automation, code reviews, and continuous integration.
  • Create and maintain technical documentation, architecture diagrams, and engineering standards. Reliability Engineering
  • Design systems with reliability, recoverability, observability, and operational excellence built in from inception.
  • Define and improve Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and operational health metrics.
  • Participate in incident response and lead the investigation of complex production issues.
  • Participate in the team's on-call rotation, providing timely triage, escalation, communication, and resolution support for production incidents impacting customer-facing applications.
  • Conduct root cause analysis and drive blameless post-incident reviews.
  • Partner with engineering teams to eliminate recurring operational issues through engineering solutions and automation.
  • Improve release safety, deployment reliability, and production readiness across services. Observability & Operational Excellence
  • Build and maintain monitoring, logging, alerting, and distributed tracing solutions.
  • Implement modern observability practices using tools such as OpenTelemetry, New Relic, Grafana, Splunk, Datadog, CloudWatch, etc.
  • Develop actionable dashboards, reliability scorecards, and service health metrics.
  • Improve detection, diagnosis, and recovery processes while reducing alert fatigue and improving signal quality. Automation Engineering
  • Eliminate operational toil through software engineering and automation.
  • Contribute to CI/CD pipelines and deployment automation. AI-Assisted Engineering
  • Leverage AI-assisted engineering capabilities to accelerate software delivery, operational workflows, incident triage, and root cause analysis.
  • Evaluate and adopt emerging AI-enabled developer productivity and reliability engineering tools. Technical Leadership
  • Mentor engineers in both software engineering and reliability engineering best practices.
  • Promote a culture of ownership, accountability, automation, and continuous improvement.
  • Influence engineering standards, platform direction, reliability objectives, and software development practices across the organization.
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.

To learn more about how we collect, keep, and process your private information, please review
Insight Global's Workforce Privacy Policy:
https://insightglobal.com/workforce-privacy-policy/. Skills and Requirements Required Qualifications
  • 7+ years of professional software engineering experience.
  • 3+ years of experience building APIs, services, or distributed systems using Node.js and TypeScript
  • 3+ years of cloud experience with AWS
  • 3+ years' experience with Splunk, New Relic and Cloud Watch
  • Experience using AI Tools
  • Experience designing and building scalable microservice architectures.
  • Experience developing, deploying, and operating production systems at scale.
  • Strong understanding of distributed systems, resiliency, scalability, observability, and performance engineering.
  • Experience with monitoring, logging, tracing, and operational tooling.
  • Experience participating in an on-call rotation or supporting production incidents for business-critical systems.
  • Experience with CI/CD pipelines, deployment automation, and modern software delivery practices.
  • Strong troubleshooting, debugging, and root cause analysis skills.
  • Excellent communication, collaboration, and leadership skills. Preferred Qualifications
  • Experience in Site Reliability Engineering, Production Engineering, Platform Engineering, DevOps, or Internal Developer Platforms.
  • Experience with Kubernetes, EKS, AKS, GKE, ECS, or other container orchestration platforms.
  • Experience implementing SLOs, SLIs, error budgets, and reliability scorecards.
  • Experience with OpenTelemetry, New Relic, Splunk, Grafana, Prometheus, Datadog, CloudWatch, Azure Monitor, or Google Cloud Operations Suite.
  • Experience supporting customer-facing web, mobile, and API platforms at scale.
  • Experience building AI-enabled applications, operational tooling, or developer productivity solutions.
  • Experience working within healthcare or other highly regulated industries.

Benefits

  • Dental Insurance