Join Team Possible!
Site Reliability Engineer (男性/女性/多样化)*
The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, performance, and scalability of both internal (in-house) and customer environments. This role acts as a critical bridge between the product platform engineering, development and customer operation team, ensuring seamless deployment, operation, and support of our solutions across many environments.
What you’ll be handling
- Maintains and optimizes availability, performance, and resilience of in-house and on-premises productive environments.
- Monitors system health, troubleshoots incidents, and implements proactive reliability solutions.
- Manages deployments, upgrades, and patching platform level changes of multiple customers.
- Demonstrates a customer-focused approach and ownership over production systems.
- Serves as technical liaison to customers (internal and external), explaining TGW’s network, software and hardware platform architecture along with delivery and deployment approach.
- Performs hands-on standing up of customer environments, and ongoing operations.
What you’ll need
- Bachelor’s degree in Computer Science, Information Technology, Computer Engineering, or a related field (or equivalent practical experience)
- Minimum of five (5) years’ experience of delivering and supporting productive IT systems following Infrastructure as Code practices.
- Up to 20% domestic and international travel.
- Hands-on experience with Linux, Kubernetes (Preferred Red Hat OpenShift) and GitOps delivery workflows.
- Experience with debugging YAML, Helm & Terraform.
- Strong troubleshooting skills across application, infrastructure, and networking layers.
What you'll receive
TGW offers full medical, dental, and vision benefits, 401K with company match, tuition reimbursement, competitive pay with PTO package offerings, along with safety shoe, protective eyewear, and fitness reimbursement.
TGW is an equal opportunity employer.
在这个职位上,您可以享受到以下福利等: