Stephen Z. Lu

prof_pic.jpg

Hi! My name is Stephen (pronounced STEH-fən) and I am a Ph.D. student in computer science at UC Berkeley supervised by Prof. Yun S. Song. I completed my undergraduate studies at McGill University in my hometown of Montreal, Canada 🇨🇦.

My research focuses on machine learning approaches for protein sequence and structure. I am broadly interested in developing deep probabilistic models of protein evolution and dynamics. Recently, my work has centered on learning representations of antibody selection during affinity maturation [1], and building generative models for protein conformation sampling [2], [3].

Previously, I dabbled in small molecule generative models [4], [5] and agentic benchmarking for autonomous scientific discovery [6].

In my free time, I love playing pickup basketball, composing some funky songs on the piano, discovering new hiking trails, and spending time with my family and friends.

Feel free to reach out at stephen.lu@berkeley.edu if you’d like to chat about research or meet for coffee in the Bay Area!

selected publications

2026

  1. ICML, 2026 · * equal contribution
    Stephen Zhewen Lu*, Aakarsh Vermani*, Kohei Sanno, and 4 more authors
    We introduce CoSiNE, a neural CTMC model of antibody affinity maturation that provably approximates the sequential point mutation process while disentangling selection from somatic hypermutation to enable inference-time affinity optimization.
    Paper

2025

  1. Aligning Protein Conformation Ensemble Generation with Physical Feedback
    ICML, 2025
    Jiarui Lu, Xiaoyin Chen, Stephen Zhewen Lu, and 4 more authors
    We introduce Energy-based Alignment (EBA), calibrating protein conformation generative models with molecular energy feedback to thermodynamically weight conformational states at state-of-the-art accuracy on MD ensemble benchmarks.
    Paper
  2. Structure Language Models for Protein Conformation Generation
    ICLR, 2025
    Jiarui Lu, Xiaoyin Chen, Stephen Zhewen Lu, and 4 more authors
    We introduce Structure Language Modeling (SLM), encoding protein structures as discrete tokens for autoregressive conformation generation that achieves a 20–100× speedup over diffusion-based methods while covering diverse ensemble modes.
    Paper