Team Introduction The team is responsible for infrastructure systems, including Storage/Computing/DB. We aim to be the leading SRE team across the industry. In the SRE team, you will have the opportunity to manage the complex challenges of scale, while using expertise in coding, algorithms, complexity analysis, and large-scale system design. We embrace a culture of diversity, intellectual curiosity, openness, and problem-solving. We also encourage ownership, self-governance and independence to work on various projects, and an environment that provides the support and mentorship needed to learn and grow as an engineer.
What you will be doing: 1. Reliability: Ensuring the reliability and efficiency of our core infrastructure, focusing on system capacity and stability; setting up reliability standards and recovery SOP. 2. Reliability: Troubleshooting and locating the technical issues, bottleneck analysis, managing system high availability architecture transformation and upgrading. 3. Efficiency: Building automated operation solutions for large-scale systems; partnering with system development teams for system iteration. 4. Efficiency: Designing and implementing software platforms and monitoring frameworks for efficient, automated, and intelligent service-oriented architecture (SOA) governance. 5. Cost: There are millions of CPUs. We should build delivery standards, and monitor and budget systems to optimize the cost of the company. 6. Compliance: Designing and setting up new IDC; designing and implementing data protection plan to meet the standard requirement.
Minimum Qualifications: - Bachelor's / Master's Degree in Computer Science or related major, with at least 5 years of relevant experience; - Solid basic knowledge of computer software, understanding of Linux operating system, storage, network IO and other related principles. - Familiar with one or more programming languages, such as Python, Go, and Java. Knowledge of design patterns and coding principles is necessary.
Preferred Qualifications: 1. Experience with storage, and relevant system experience with the following: KV, Table, Graph, Redis, MySQL, MongoDB, MQ, and Kafka. 2. Experience with computing & big data, and system experience with the following: Kubernetes, Docker/Containers, AIops, Spark, Flink, Function as a service, RPC Framework, and Service Mesh.
Seen 21 days ago.
Original posting on ByteDance's site ↗
Posting text belongs to the employer. Removal requests: contact us.
Live postings like this one
Financial Service Partnership Director, AMS - Global Payment
ByteDance
San Jose, California, United States of America
9h agoRisk Manager - Global Payment - Singapore
ByteDance
Singapore
22h agoAI & Cloud Sales Manager - BytePlus
ByteDance
Hong Kong (China), Hong Kong Island, Hong Kong, China
2d agoEnterprise Account Executive (New Logo) - Lark APAC, HCMC
ByteDance
Ho Chi Minh City, Ho Chi Minh, Vietnam
2d agoSenior Product Manager - Corporate Information Systems - AI Vertical Team
ByteDance
Dubai, United Arab Emirates
2d agoEnterprise Account Executive(New Logo )- Lark APAC, Hanoi
ByteDance
Hanoi, Ha Noi, Vietnam
2d agoSenior Software Engineer - Search & Vector Database Infrastructure
ByteDance
Seattle, Washington, United States of America
2d agoGlobal Strategic Partnerships Director (Philippines E-Commerce)
ByteDance
Taguig, National Capital Region (Manila), Philippines
3d ago