Advertisement
नौकरियों पर वापस जाएँ

Site Reliability Expert

Valtech

Canadaपूर्णकालिकरिमोटCAD 100,000 – 150,000आवेदन की अंतिम तिथि 30 अक्टू॰ 2026

भूमिका के बारे में

Why Valtech ? We’retheexperience innovation company - a trusted partner to the world’s most recognized brands. To our people we offer growth opportunities, a values-driven culture, international careers and the chance to shape the future of experience.

The opportunity At Valtech , you’ll find an environment designed for continuous learning, meaningful impact, and professional growth. Whether you're pioneering new digital solutions, challenging conventional thinking or building the next generation of customer experiences, your work will help transform industries. We are proud of: The work we do and the innovation we drive Our values of share, care and dare A workplace culture that fosters creativity, diversity and autonomy Our borderless, global framework, which enables seamless collaboration The role Please be aware thet French speaking skills are needed for this role.

We are seeking a highly experienced Site Reliability Expert to lead and drive observability, reliability, and operational excellence initiatives across complex cloud-native environments. This role goes beyond platform administration and requires a strong Site Reliability Engineering (SRE) background, combining observability expertise with production operations, automation, and reliability best practices. The ideal candidate will be an experienced technical leader capable of defining observability standards and strategies, supporting product teams, implementing reliability practices, and enabling scalable monitoring solutions across distributed microservices architectures.

You will thrive in this role if you are: A curious problem solver who challenges the status quo A collaborator who values teamwork and knowledge-sharing Excited by the intersection of technology, creativity and data Experienced in Agile methodologies and consulting (a plus) Role

responsibilities

Define and implement observability strategies, standards, and governance across applications and platforms. Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms. Establish and drive SRE best practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.

Partner with engineering and product teams to improve system reliability, performance, and operational maturity. Develop standards for tagging, ownership, dashboard design, access management, and alerting governance. Support teams that are not specialized in observability by providing guidance, coaching, and knowledge transfer.

Lead technical workstreams, prioritize initiatives, and ensure successful delivery within defined timelines and budgets. Analyze distributed systems and troubleshoot complex production issues using monitoring and tracing data. Promote documentation, operational rigor, and continuous improvement across engineering teams.

Collaborate effectively within a distributed, multilingual environment. Must have qualifications To be considered for this role, you must meet the following essential qualifications: Site Reliability Engineering & Observability Significant experience in Site Reliability Engineering (SRE) within large-scale production environments. Deep understanding of: Service Level Indicators (SLIs) Service Level Objectives (SLOs) Error Budgets Symptom-based Alerting Proven expertise with enterprise observability platforms such as: Dynatrace Datadog New Relic AppDynamics Strong experience with: Application Performance Monitoring (APM) Real User Monitoring (RUM) Monitoring agents and instrumentation Alerting strategies Role-Based Access Control (RBAC) SLO management Tagging and governance models Distributed Systems & Cloud Platforms Strong knowledge of OpenTelemetry (OTEL) and distributed tracing.

Experience working within composable, microservices-based architectures. Hands-on production experience with: AWS Kubernetes Automation & DevOps Experience with infrastructure and operational automation. Practical knowledge of: Terraform Bash scripting Python scripting Experience with CI/CD tools such as GitLab CI or equivalent pipeline/workflow platforms.

Delivery & Collaboration Demonstrated ability to lead technical initiatives and workstreams. Experience working within complex operational and agile environments. Strong stakeholder management and collaboration skills.

Excellent communication skills in both French and English. Strong documentation practices, organizational skills, and attention to detail. High degree of autonomy and ownership.

Nice to have qualifications Experience monitoring and supporting Java Spring Boot applications. Experience within e-commerce platforms and high-transaction environments. Experience establishing enterprise-wide observability frameworks and governance models.

Consulting or advisory experience supporting multiple engineering teams. If you do not meet all the listed qualifications or have gaps in your experience, we still encourage you to apply. At Valtech , we recognize that talent comes in many forms, and we value diverse perspectives and a willingness to learn.

The

benefits

This is a Full time position based in Canada . The offered salary range is $100,000 - $150,000 CAD annually, depending on experience and location. Valtech offers a comprehensive benefits package effective after three months of continuous service: A comprehensive insurance plan , where you can choose the module that best suits your needs—Gold, Silver, or Bronze.

The employer may contribute up to 80% of your coverage depending on the selected module. This plan includes short- and long-term disability coverage . Dialogue via Sun Life provides virtual healthcare services, allowing you to consult with a healthcare professional for emergencies, prescription renewals, and more.

You also have access to the Employee and Family Assistance Program , as well as a complete mental health support program . A $500 Personal Spending Account , which can be used

आवश्यक कौशल

Site Reliability Engineering
SRE Engineering
DevOps Engineer
Platform Engineering
Cloud Engineer
Site Reliability Engineer
Site Reliability Operations Engineer
Principal Site Reliability Engineer
Site Reliability Engineering (SRE)
Site Reliability Engineering Lead
python
aws
kubernetes
english_language
Advertisement