← All jobs

Senior Database Reliability Engineer

About the role

• Ensure high availability, performance, and reliability of MySQL and PostgreSQL databases as critical services from an SRE/DBRE perspective
• Define, implement, and continuously improve SLOs, SLIs, and SLAs, monitoring error budgets and taking proactive action
• Design and operate scalable, resilient cloud architectures, including Multi-AZ, read replicas, sharding, and partitioning
• Perform query, index, execution plan, parameter, connection, and connection-pooling tuning and optimization in partnership with development teams
• Define and maintain backup, restore, point-in-time recovery, and disaster recovery strategies
• Automate operational tasks using IaC and scripting, reducing manual processes and toil
• Plan for capacity and data growth while optimizing cloud resources and costs
• Implement observability through metrics, logs, traces, dashboards, and alerts
• Plan and execute engine and workload migrations and upgrades with minimal impact
• Define and promote standards, guidelines, and frameworks for data modeling, access, and operations
• Assess and resolve critical incidents, leading root cause analyses and preventive actions
• Ensure security and compliance, including encryption, access controls, secrets management, data masking, LGPD, and PCI DSS
• Conduct design reviews and schema reviews
• Promote operational excellence, reliability, and data engineering best practices
• Evaluate new AWS data technologies and services

• Proven professional experience administering, operating, and tuning relational databases in large-scale production environments
• Deep knowledge of MySQL, including replication, storage engines, query optimization, parameters, and troubleshooting
• Deep knowledge of PostgreSQL, including MVCC, vacuum/autovacuum, indexes, execution plans, extensions, and replication
• Experience with AWS and data, networking, and security services: RDS, Aurora, DMS, S3, IAM, VPC, KMS, Secrets Manager, and Parameter Store
• Experience applying SLOs/SLIs, error budgets, observability, and toil reduction to data services
• Experience tuning and optimizing performance, including queries, indexes, locks, contention, and connection pooling
• Experience with PgBouncer, ProxySQL, or RDS Proxy
• Experience with high availability, replication, failover, backup, restore, and disaster recovery
• Experience with IaC and automation using Terraform or AWS CDK (Python)
• Shell/Bash scripting skills
• Experience with Grafana, Prometheus, Datadog, New Relic, CloudWatch, or Performance Insights
• Knowledge of data security and compliance, including encryption, access management, secrets management, and auditing
• Knowledge of CI/CD and schema versioning/migrations, such as Flyway or Liquibase
• Ability to create and maintain technical documentation, runbooks, and diagrams
• Preferred qualifications: AWS certifications; experience in fintech or regulated environments; large-scale migrations; NoSQL/in-memory databases; streaming; FinOps; containers; Chaos Engineering; advanced observability; networking; Git/GitHub/GitFlow; and Python, Go, Java, or Node.js

• Medical and dental insurance with no co-pay
• Life insurance
• Prescription medication allowance
• Fitness allowance
• Four free therapy or nutritionist sessions per month
• Quick massage at headquarters
• Flexible meal benefit on a Visa card
• Complimentary food at headquarters
• Childcare allowance
• Parental support program
• Extended maternity and paternity leave
• In-house training platform
• Education allowance covering 70% of undergraduate and language course tuition, as well as courses and books
• Home office allowance
• Work equipment
• Furniture allowance
• Partnership with WOBA for coworking access throughout Brazil
• Birthday month day off
• Happy hour allowance
• Referral bonus for new hires
• Annual performance-based bonus
• Stock options plan
• No dress code