Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details
Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

TEKGENCE INC

Lead Cloud Platform & Infrastructure Engineer

Career Insights for Platform Engineer

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on California data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

A Platform Engineer is responsible for the development of platforms that support the needs and use cases of different engineering teams across the organization. Creates reusable tools and workflows to streamline operational needs and facilitate automation tasks, supporting scalability of DevOps practices.

$161,754 / year median in California

Explore Career

Job Description

Lead Cloud Platform & Infrastructure Engineer at

TEKGENCE INC

Lead Cloud Platform & Infrastructure Engineer at

TEKGENCE INC

in Los Altos, California Posted in about 11 hours ago.

Type:

full-time

Job Description:

Core Platform Engineer (L1, breadth-first) Sunnyvale, CA or San Jose, CA (5 days Onsite, Final round F2F) Contract & FTE (Both)

Job Description:

- The engineering first line of defense: incident response, triage, reliability, and automation across the full infrastructure stack.

Day-to-day:

Runs incident response drills, post-mortems, and root cause analysis; learns from past incidents to prevent recurrence. Starts the day reviewing overnight alerts and system performance metrics, triaging anomalies. Participates in team stand-ups on projects, incidents, and daily priorities. Automates routine processes, analyzes system logs, and builds tools to strengthen monitoring. Works alongside software engineers advising on resilient-code best practices and reviewing changes pre-deployment. Maintains high SLIs/SLOs; documents work and shares insights with a customer-centric mindset.

Must have:

Architecture, design patterns, reliability, and scaling of new and existing systems. Incident command experience - driving RCA, coordinating cross-functional teams, ensuring corrective-action follow-through. Observability built from the ground up - defining SLOs/SLIs, closing monitoring gaps, alerting strategies that catch failures before customers do. Linux kernel internals - scheduler, memory allocation, driver subsystems. High-quality code in at least one language (Python, Go, or similar). System-level debugging - kdump, kernel panic analysis. IaC (Ansible, Terraform, Kubernetes) and CI/CD (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure. TCP/IP and network programming. Distributed storage systems - object, block, and/or file storage paradigms. Strong communication skills.

Nice to have:

Hardware and GPU troubleshooting. OVN/OVS-based networking stack exposure.

Direct:

469-421-5604 , Ext- 218

• nitesh.j@tekgence.com

Linkedin:

linkedin.com/in/nitesh-ch-a378b5222 6655 Deseo Dr, Suite 104,Irving, TX , 75039

• www.tekgence.com