Machine Learning Scientist - Generative AI for Protein Design
发布于: 30/07/2026
Shanghai East China
Permanent
生命科学、医疗保健、制药
We are seeking a creative and rigorous Machine Learning Scientist to develop next-generation generative models for protein and biologics design. The scientist will conduct original research at the intersection of generative AI, geometric deep learning, protein language modeling, and structure-based molecular design, with particular emphasis on diffusion or flow-based models, large protein language models, and E(3)/SE(3)-equivariant architectures.
The successful candidate will formulate biologically meaningful design problems, develop and train new models across protein sequence and three-dimensional structure, and translate model outputs into experimentally testable hypotheses. Working in close partnership with antibody/protein engineers, structural biologists, biophysicists, data scientists, and ML engineers, this individual will help establish closed-loop design-make-test-learn workflows for biologics discovery.
Essential Responsibilities:
• Develop novel generative ML methods for protein sequence, backbone, side-chain, complex, and interface design using diffusion, score-based, flow-matching, autoregressive, masked, or hybrid approaches.
• Design and implement E(3)/SE(3)-equivariant or invariant neural architectures that respect biomolecular geometry, including graph neural networks, geometric transformers, frame-based models, and related methods.
• Build and adapt protein language models and multimodal foundation models that jointly reason over sequence, structure, function, text, assay, and property data.
• Create conditional generation strategies for target-aware binder design, motif scaffolding, inverse folding, interface optimization, antibody or protein engineering, and multi-parameter property optimization.
• Integrate biological and physical constraints into generation, including symmetry, residue or motif constraints, sterics, geometry, energetics, developability, and manufacturability considerations.
• Develop robust training datasets and data pipelines from structural, sequence, assay, and literature-derived sources; establish sound split strategies that minimize leakage and accurately assess generalization and novelty.
• Define fit-for-purpose evaluation frameworks covering structural plausibility, sequence recovery, diversity, novelty, target interaction, uncertainty, computational efficiency, and prospective experimental success.
• Use structure prediction, inverse folding, docking, molecular modeling, and physics-based tools to triage and refine generated candidates while recognizing the limitations of in silico scoring.
• Partner with experimental scientists to select candidates, interpret assay results, learn from failures, and iteratively improve models through design-make-test-learn cycles or active learning.
• Conduct ablation studies, benchmarking, error analysis, and mechanistic investigation to understand model behavior and establish scientific credibility.
• Publish high-impact research, contribute to patents where appropriate, and communicate methods and results clearly to scientific and project teams.
• Write reproducible, well-tested research code and collaborate with ML engineering and MLOps colleagues to scale training, inference, model tracking, and deployment.
• Monitor advances in generative modeling and computational protein design; translate promising ideas into practical capabilities for biologics programs.
Qualifications Required:
• Master's degree or Ph.D. with research or industry experience in machine learning, computational protein science, structural bioinformatics, geometric deep learning, generative modeling, or a closely related area. Candidates with a Master's degree should typically bring substantial hands-on research or industry experience and a demonstrated record of technical contributions.
• Demonstrated hands-on experience developing generative models, with depth in at least two of the following: diffusion/score-based models, flow matching, autoregressive or masked language models, variational methods, energy-based models, or reinforcement learning for molecular design.
• Strong experience with geometric deep learning and three-dimensional molecular representations, including E(3)/SE(3)-equivariant models, graph neural networks, geometric transformers, local frames, rotations, or manifold-aware learning.
• Experience with protein language models or biological foundation models and an understanding of protein sequence-structure-function relationships.
• Strong knowledge of modern deep-learning methods, including transformers, attention, representation learning, conditional generation, sampling, fine-tuning, and uncertainty or calibration.
• Proficiency in Python and PyTorch or JAX, with the ability to independently implement, train, debug, and evaluate research models on multi-GPU systems.
• Experience working with protein sequence and structure data and common formats or resources; practical familiarity with structure prediction and protein-design workflows.
• A rigorous approach to dataset construction, leakage control, benchmark design, statistical analysis, reproducibility, and interpretation of negative results.
• A strong research record demonstrated through first-author or substantial-contribution publications at leading ML or computational biology venues, impactful open-source work, patents, or equivalent evidence.
• Excellent scientific communication and collaboration skills, including the ability to work across ML, computational, and experimental disciplines.
• Professional working proficiency in English; Chinese communication capability is strongly preferred for the Shanghai-based team.
Preferred Qualifications:
• Direct experience with protein or antibody generation systems such as RFdiffusion-family methods, Chroma, ProteinMPNN/LigandMPNN, ESM-family models, ProGen, diffusion protein language models, or comparable internal methods.
• Experience designing binders, antibodies, enzymes, peptides, protein complexes, or constrained functional motifs, ideally with prospective experimental validation.
• Familiarity with antibody-specific challenges such as CDR representation, antigen conditioning, affinity/specificity optimization, human-likeness, immunogenicity, aggregation, stability, and other developability properties.
• Experience with multimodal or all-atom modeling, co-design of sequence and structure, preference optimization, active learning, Bayesian optimization, or lab-in-the-loop learning.
• Knowledge of molecular modeling or structural bioinformatics tools such as AlphaFold, Rosetta, OpenFold, molecular docking, or molecular dynamics.
• Experience training large models with distributed computing and mixed precision; familiarity with experiment tracking and reproducible ML workflows.
• AWS cloud-based ML experience is an asset, particularly GPU training, S3 data workflows, containers, SageMaker, AWS Batch, EKS, or related services.