All jobs
Visa logo
Visa

200 open roles

Staff Site Reliability Engineer

Full-time
Lead
Posted 2 days ago
Apply on company siteYou’ll apply directly with the employer on Workday.
Added to ZestAmigo
Last seen on employer site

Where you can work

Remote

Open to candidates in Brazil

View location wording from the posting
BR - Remote - Brazil, Brazil

Employer description

About Us

Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid.

At Visa, you'll have the opportunity to create impact at scale — tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world.

Join Visa and do work that matters – to you, to your community, and to the world. Progress starts with you.

Job Description

The Staff Site Reliability Engineer is responsible for architecting, implementing, and maintaining solutions that ensure Visa’s application services operate with high availability and reliability. This role contributes by helping define the reliability expectations that services must meet, guiding application teams through resilient architecture and design choices, and creating repeatable evidence that critical services can tolerate realistic failure modes. Because the team is still small, the candidate must be able to work independently, lead complex technical work with limited guidance, and influence application-owning teams without direct authority.

All roles require digital fluency, including the ability to work with emerging technologies such as Generative AI tools (e.g. ChatGPT, Microsoft Copilot) to support everyday work.

The Reliability & Resilience Engineering squad works to improve the reliability and resilience of the services sold to customers. The team provides subject matter expertise in chaos engineering, reliability standards, resilience design, observability, and continuous reliability improvement. This consultant-level role is expected to operate as a high-autonomy technical owner within the function, creating clarity from ambiguity and translating business reliability goals into practical engineering standards, experiments, technical guidance, and improvement plans.

Key Responsibilities:

  • Own complex resilience engineering work across priority services.
  • Design, plan, conduct, and report on controlled chaos experiments and game days.
  • Establish repeatable chaos testing patterns.
  • Define reliability and resilience standards.
  • Translate incidents, observed failure modes, SLO misses, and experiment findings into actionable engineering improvements.
  • Provide architecture and system-design guidance to application teams, especially around circuit breakers, load shedding, timeouts, retries, dependency isolation, graceful degradation, observability, alerting, and failure containment.
  • Author technical design documents, experiment plans, standards, post-experiment reports, and improvement proposals.
  • Mentor engineers through design reviews, technical guidance, and hands-on support to raise the reliability capability of the teams they work with.
  • Deliver documented resilience standards, reusable experiment templates, clear production readiness criteria, completed pilot experiments, prioritised remediation backlogs, and measurable improvements in the resilience posture of selected customer-critical services.

This is a remote position. A remote position does not require job duties be performed within proximity of a Visa office location. Remote positions may be required to be present at a Visa office with scheduled notice. #LI-Remote

Qualifications

Basic Qualifications:

  • 5+ years of relevant work experience with a Bachelor’s Degree or at least 2 years of work experience with an Advanced degree (e.g. Masters, MBA, JD, MD) or 0 years of work experience with a PhD, OR 8+ years of relevant work experience.
  • Experience in Kubernetes and related technologies such as Helm, ArgoCD, and Terraform.
  • Experience in automating complex tasks and processes using programming languages such as Python, Go, and Java.
  • Experience in planning and performing chaos experiments using tooling such as Gremlin, Chaos Toolkit, Litmus, or equivalents

##

Preferred Qualifications:

  • 6 or more years of work experience with a Bachelor's Degree or 4 or more years of relevant experience with an Advanced Degree (e.g. Masters, MBA, JD, MD) or up to 3 years of relevant experience with a PhD.
  • Experience working with Agile teams.
  • Experience in fast-paced 24x7 environments.
  • Experience in distributed systems design and maintenance.
  • Experience in cloud and system architecture across hybrid platforms (AWS, GCP, Azure).
  • Experience in observability and performance optimization techniques.
  • Experience working as a software developer

Visa is an EEO Employer

Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability or protected veteran status. Visa will also consider for employment qualified applicants with criminal histories in a manner consistent with EEOC guidelines and applicable local law.

Track this application

Keep your own notes. Only you can mark an application as sent.

Report a problem with this listing

Sign in to report this listing.

More at Visa

Similar roles elsewhere