Full SRE Analyst - Affirmative Vacancy for Women

Brazil | Sept. 14, 2026

Report as Closed

Company: Experian

Country: Brazil

Type: Remote

Employment: Full-time

Description:

We are looking for a highly motivated Site Reliability Engineer (SRE)to join our Cloud, Data & AI Platform team. In this role, you will be responsible for designing, operating, and continually improving the reliability, scalability, observability, and performance of cloud-native platforms that support business-critical applications, data pipelines, and AI/ML workloads.

You will work closely with the Software Engineering, Data Engineering, AI Engineering, and Platform teams to build resilient systems, automate operations, improve the developer experience, and establish reliability best practices across the organization.

As an Engineer fully, you will play a key role in advancing operational excellence through automation, observability, incident management, cost optimization and infrastructure modernization.

Key Responsibilities:

  • Design and operate highly available, scalable, and secure cloud platforms on AWS.
  • Build and maintain Kubernetes-based infrastructure to support applications, data, and AI workloads.
  • Improve platform reliability through automation, Infrastructure as Code (IaC), and self-service capabilities.
  • Implement and improve observability solutions using Datadog, including monitoring, logs, tracking, alerts, dashboards and SLO management.
  • Support and optimize large-scale data processing environments using Airflow, Amazon EMR, S3, and other AWS data services.
  • Work in partnership with Data and AI teams to improve the reliability, scalability and operational maturity of Machine Learning and Artificial Intelligence platforms.
  • Lead incident response activities, root cause analyzes and post-incident reviews, promoting continuous improvement.
  • Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets.
  • Improve deployment processes, CI/CD pipelines, and release reliability.
  • Optimize the utilization, performance and costs of cloud infrastructure.
  • Mentor team members and promote SRE best practices across the engineering organization.

What defines success in this role

  • Increased platform availability and reliability.
  • Improved observability and reduced incident resolution time.
  • Greater automation and reduction of manual and repetitive operational efforts.
  • Reliable, scalable and cost-efficient Data and AI Platforms.
  • Strong collaboration with Engineering teams to deliver resilient production systems.

Mandatory qualifications

  • Higher education currently attended/completed.
  • Solid experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering or DevOps roles.
  • Strong practical experience with AWS services and cloud-native architectures.
  • In-depth knowledge of Kubernetes and containerized workloads in production environments.
  • Experience in managing and solving problems in large-scale distributed systems.
  • Solid experience with observability platforms, preferably Datadog.
  • Experience in supporting platforms and data flows using technologies such as Airflow, EMR, Spark and S3.
  • Expertise in Infrastructure as Code (IaC) using Terraform or similar tools.
  • Experience in building and maintaining CI/CD pipelines and platform automation.
  • Solid knowledge of Linux, networking and system performance troubleshooting.
  • Proficiency in scripting and automation using Python, Bash or similar languages.

Desirable qualifications

  • Experience in supporting large-scale cloud-native platforms in AWS environments.
  • Experience with Kubernetes platform operations and cluster lifecycle management.
  • Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, error budgets and operational excellence practices.
  • Experience in implementing observability solutions using tools such as Datadog, Prometheus, Grafana, OpenTelemetry or similar technologies.
  • Familiarity with data processing and workflow orchestration platforms, such as Airflow, Spark or EMR.
  • Experience with Infrastructure as Code (IaC) and platform automation practices.
  • AWS, Kubernetes, Terraform or Datadog certifications.
  • Experience working in large-scale, highly available corporate environments

    Apply here:

    Web: Apply here

    Emails:



Similar Python Jobs

Found 6 similar Remote jobs

Principal Software Engineer - Python

Lifelancer

Remote Full-time

United States

$150,000 - $300,000

View Job →
MySQL Database Administrator

Miratech

Remote Full-time

India

View Job →
Sr Data Integration Engineer

HomeServices of America

Remote Full-time

United States

$153,000 - $178,000

View Job →
Growth Automation Engineer

The Global Talent Co.

Remote Full-time

Brazil

View Job →
AI Prompt Designer

Bright Vision Technologies

Remote Full-time

United States

$100,000 - $100,000

View Job →
Staff Software Engineer, Security Engineering

Coalition

Remote Full-time

United States

$160,000 - $230,000

View Job →

Find High-Paying Python Developer Jobs ($80K-$200K+)

Django · Flask · FastAPI · Remote & Onsite · Updated daily · No recruiter spam
💼 Get the best Python jobs weekly. Salary-transparent roles, no recruiter spam — unsubscribe anytime.