Job Posting Organization: The Wikimedia Foundation is a nonprofit organization that operates Wikipedia and other Wikimedia free knowledge projects. Established with the vision of a world where every individual can freely share in the sum of all knowledge, the foundation believes in the potential of every person to contribute to shared knowledge and the importance of free access to that knowledge. The organization is supported by donations from millions of individuals globally, averaging around $15 per donation, alongside institutional grants. As a 501(c)(3) tax-exempt organization in the United States, it has its headquarters in San Francisco, California, and operates with a diverse workforce across more than 40 countries. The foundation is committed to maintaining an inclusive and equitable workplace, encouraging applications from individuals with diverse backgrounds.
Job Overview: The Senior Site Reliability Engineer at the Wikimedia Foundation will be integral in designing, developing, and maintaining a reliable, scalable, and highly available infrastructure for the organization's API services. This role involves significant contributions to the challenges of innovating and maintaining Wikipedia's data feeds for high-volume users. The engineer will collaborate across departments, particularly with the Wikimedia Foundation's SRE teams, to meet reliability targets for critical APIs while balancing performance, cost, and availability through data-driven decisions. The position also entails designing and managing infrastructure and services that support Wikimedia's projects, including Kubernetes clusters and application servers, and participating in incident response and on-call duties. The role is dynamic and fast-paced, requiring a proactive approach to building services that enhance the distribution of knowledge globally.
Duties and Responsibilities: The responsibilities of the Senior Site Reliability Engineer include defining, tracking, and improving Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to ensure reliability targets are met. The engineer will build and enhance observability systems for proactive detection and troubleshooting, drive reliability engineering practices such as capacity planning and load testing, and improve developer experience by enabling self-service infrastructure. They will partner with engineering teams to embed reliability best practices early in the development lifecycle, design and optimize CI/CD and GitOps workflows, implement secure infrastructure practices, and continuously optimize infrastructure cost and efficiency. The engineer will also establish operational metrics to drive continuous improvement, reduce operational toil through automation, and contribute to evolving internal platform capabilities. Mentoring peers in technical and operational areas is also a key responsibility.
Required Qualifications: Candidates for the Senior Site Reliability Engineer position should possess experience with Infrastructure as Code and automation tools, proficiency in at least one programming language, and experience designing and optimizing cloud-based systems. Familiarity with CI/CD pipelines and GitOps workflows is essential, along with experience in incident management and reliability operations. A strong understanding of SRE principles, including observability practices, is required, as well as the ability to work effectively in a distributed, cross-functional environment with strong documentation and communication skills. Familiarity with Wikimedia or other open-source projects is considered a plus.
Educational Background: While the job description does not specify exact educational requirements, a background in computer science, engineering, or a related field is typically expected for such technical roles. Relevant certifications in cloud infrastructure or site reliability engineering may also be beneficial.
Experience: The position requires proven experience operating highly available, large-scale distributed systems, with a deep understanding of reliability, scalability, and failure modes. Candidates should demonstrate an ownership mindset, taking responsibility for system reliability and proactively identifying risks. A bias for automation and a continuous improvement mindset are also important, as is a focus on customer experience and adaptability in a fast-evolving environment.
Languages: While specific language requirements are not mentioned, proficiency in English is likely mandatory due to the global nature of the team and the need for effective communication. Additional language skills may be advantageous, particularly in regions where Wikimedia operates.
Additional Notes: The Wikimedia Foundation is a remote-first organization, allowing staff members to work from various locations, including over 40 countries. The anticipated annual pay range for this position within the United States is between $116,633 and $181,243, with adjustments based on location for international applicants. The organization values competitive and equitable salaries, and does not consider salary history in its hiring process. The role is full-time, and the foundation encourages applications from individuals requiring accommodations due to disabilities.
Info
Job Posting Disclaimer
This job posting is provided for informational purposes only. The accuracy of the job description, qualifications, and other details mentioned is the sole responsibility of the employer or the organization listing the job. We do not guarantee the validity or legitimacy of this job posting. Candidates are advised to conduct their own due diligence and verify the details directly with the employer before applying.
We are not liable for any decisions or actions taken by applicants in response to this job listing. By applying, you agree that all application processes, interviews, and potential job offers are managed exclusively by the listed employer or organization.
Beware of fraudulent job offers. Do not provide sensitive personal information or make any payments to secure a job.