The Technical Infrastructure SRE team is responsible for managing the whole infrastructure and applications. Our mission is to ensure all production systems can support our fast growing world-wide user base as well as keep the entire systems stable, efficient and cost effective. We manage deployments, system capacity, traffic scheduling, fault tolerance, disaster recovery, emergency response, automations, operation platforms development, etc.
Be responsible for the basic engineering construction of byte infrastructure products & components, focusing on infrastructure O&M architecture optimization, automated O&M platform research and development, data and intelligent O&M. Through the methodology of software engineering and digital intelligence, O&M, around the O&M requirements of infrastructure products & components, built a layered and systematic O&M platform to solve the problem of ultra-large-scale cluster O&M management. (Goals) To provide stable, efficient, and low-cost serverless infrastructure facilities for Mid-Platform & Business.We aim to be the leading SRE team across the industry.
- Grow and lead a team of engineers committed to building and operating scalable and reliable Infrastructure Platform systems.
- Be both technically hands-on and people manager.
- Provide technical leadership and guidance to both your team members and your project peers.
- Communicate cross-functionally across various teams, organizations and internal and external stakeholders to drive engineering efforts.
- Lead the team's innovation efforts, bring in new ideas and technologies.
Minimum Qualifications 1. Expertise in analyzing and troubleshooting distributed systems. 2. Bachelor/Master's degree in Computer Science, a related technical field involving software development or systems engineering. 3. Experience programming in at least one of the following languages: Python or Golang
Preferred Qualifications 1. Excellent communication skills and ability to collaborate cross-functionally with data science and infrastructure teams. 2. Hands-on in designing, building, scaling, and troubleshooting platform solutions. 3. Strong understanding of code optimizing and routine tasks automation. 4. Experience with compute/storage/database, and relevant system experience with the following: HDFS, Object storage, file storage, KV, Table, Graph, Redis, MySQL, MongoDB, MQ, and Kafka. Kubernetes, Docker/Containers, AIops, Spark, Flink, Function as a service, RPC Framework, and Service Mesh.
Seen 2 hours ago.
Original posting on TikTok's site ↗
Posting text belongs to the employer. Removal requests: contact us.
Nearby
Live postings like this one
Same employer first, then the same role elsewhere.
- 2h ago
- 2h ago
Home Improvement Account Manager Project Intern (TikTok Shop by Tokopedia) - 2026 Start
Jakarta, Jakarta Raya, Indonesia
2h agoTikTok Commerce - Growth Strategy Manager, Business Planning and Operation
Istanbul, İstanbul, Turkey
2h ago- 2h ago
- 2h ago
Technical Sourcer, Product Management - San Jose (Third-party Associate)
San Jose, California, United States of America
2h ago- 2h ago
One job at a time
One posting. One CV. $25.
Pick the job you actually want and we write for it.