The posting
Job Responsibilities
- Pre-training Strategy & Architecture
- Pre-training Data Engineering
- Large-scale Distributed Training
- Long-context Training
- Training Monitoring & Optimization
- Evaluation & Iterative Optimization
Job Requirements
- Bachelor's degree or above in Computer Science, Artificial Intelligence, Natural Language Processing (NLP), Machine Learning, Distributed Systems, or a related technical field.
- Hands-on experience delivering or contributing to end-to-end LLM pre-training projects.
- Proven experience participating in the pre-training of models with 7B+ parameters.
- Strong hands-on expertise with distributed training frameworks such as Megatron-LM, DeepSpeed, or FSDP.
- Practical experience working with large-scale training environments involving 64+ GPUs.
- Strong understanding of large-scale pre-training data pipelines, including data cleaning, deduplication, quality filtering, tokenization, data mixing, and data quality optimization.
- Experience designing or optimizing data pipelines for large-scale LLM training.
- Strong ability to analyze training loss, gradients, convergence, and training stability.
- Proven experience troubleshooting and resolving issues in large-scale distributed training environments.
- Familiarity with long-context training and context-extension techniques, including
- RoPE scaling, NTK-aware interpolation, and YaRN.
Preferred Qualifications
- Experience with 70B+ parameter model pre-training.
- Experience with Mixture-of-Experts (MoE) model pre-training.
- Experience optimizing large-scale GPU clusters, distributed training systems, and AI training infrastructure.
- Publications in top-tier AI/ML conferences such as NeurIPS, ICML, ICLR, ACL, or EMNLP, particularly in areas related to LLM pre-training, model architecture, or training optimization.



