Site Reliability Engineer
Velozient
Job Description
We are looking for a full-time, remote Senior Site Reliability Engineer with 3+ years of experience in high-availability cloud environments using AWS to join our U.S. client. This role is heavily focused on AWS infrastructure, CloudFormation, and CI/CD pipeline ownership, with an emphasis on improving reliability, performance, and cost efficiency.
This position is ideal for someone who operates with an engineering mindset, takes strong ownership, and is comfortable recommending and implementing improvements across infrastructure and pipelines. You will join a shared SRE/DevOps function, working closely with engineering teams to ensure scalable, secure, and cost-effective systems. The role also requires readiness to work within and manage Windows-based environments, as the client operates primarily on Windows infrastructure.
Our client provides industry-leading environmental compliance, fuel management, and facility maintenance software used by major fuel networks, including global brands. Their platform helps companies monitor and optimize operations by delivering real-time insights, automating compliance tasks, and ensuring site uptime.
With a mission to make sustainability and operational efficiency seamless, our client is transforming how fuel networks operate. As part of their dynamic, fast-paced engineering team, you'll have the opportunity to work on cloud-based, data-driven solutions that have a direct impact on performance, compliance, and environmental responsibility.
Responsibilities
- Own and improve AWS-based infrastructure, with a strong focus on CloudFormation templates
- Design, implement, and optimize CI/CD pipelines using Azure DevOps and AWS integrations
- Analyze and manage infrastructure costs, proactively identifying optimization opportunities
- Work closely with engineering teams to reduce their dependency on writing CloudFormation templates, improve deployment processes and operational efficiency, and recommend and implement improvements across existing and new pipelines
- Ensure proper monitoring, observability, and reliability standards before production releases
- Support release processes, including availability during release windows when needed
- Build and maintain reusable infrastructure and automation modules
- Contribute to security best practices within cloud infrastructure
- Collaborate in a safe Agile environment, participating in sprints and technical discussions
Required Experience
- Excellent English communication skills
- 3+ years supporting a high-availability cloud environment in a DevOps, SRE, or Sysadmin role
- Strong, hands-on experience with AWS and Azure
- Experience building and managing CI/CD pipelines, particularly using Azure DevOps and AWS-native tools or integrations
- Systems and infrastructure deployment and management practices
- Readiness to work in and manage Windows-based environments
- AWS Infrastructure & CloudFormation – Strong hands-on experience managing production AWS environments, with expertise in CloudFormation, networking, and high-availability infrastructure
- Windows Server Administration – Experience administering and maintaining Windows Server environments, including patch management, troubleshooting, and production support
- Windows Workloads in AWS/Azure – Proven experience deploying, managing, and supporting Windows-based workloads across AWS and/or Azure cloud environments
- PowerShell & Automation – Strong PowerShell scripting skills for automation, infrastructure management, and operational tasks
- CI/CD & SRE Mindset – Experience with Azure DevOps pipelines, infrastructure automation, monitoring, incident response, and improving system reliability and operational efficiency
- Proven delivery experience in a vibrant, dynamic startup environment
- A collaborative approach, a can-do attitude, and a relentless pursuit to attain goals and solve problems
- Trustworthy, team-oriented, and transparent
Desired Experience
- University degree or relevant industry experience
- Extensive networking skills, including VPC, VPN, TCP/IP, routing, load balancing, or DNS
- Experience provisioning, deploying, and supporting containers with Helm, Kubernetes (Kops), or Docker
- Scripting and development experience using Python, Go, or bash with associated source control
- Hands-on monitoring tool usage, including AWS CloudWatch, Datadog, or similar tools
- Experience with relational databases such as MySQL and PostgreSQL, and document stores such as MongoDB
- Microsoft Azure Cloud experience
Additional Information
- Knowing your ideas are heard and matter think big!
- You get to own your job and be recognized for your contributions
- Work with smart and creative people
- Making mistakes is human. Let's learn from them. Be transparent!
- We recognize you as an individual no presumptions or judgment. Be the extraordinary you!
- 15 days Paid Time Off (PTO), 1 floating day, 3 sick days, and designated national holidays
- Start: ASAP
About Velozient
We are a privately held, nearshore software development company providing outsourced development resources to North American companies. Our mission is to offer development talent that enjoys taking on challenging work, wants to grow their skills and experiences building software, and excel in a fast-paced, dynamic team environment. We are focused on providing world-class remote resources to work as valued client team members.
If this type of opportunity excites you, then consider joining our team!