Yang Li (郦洋)
Ph.D. Student @ ReThinkLab, School of AI, SJTU.
About
I am a Ph.D. student at the School of Artificial Intelligence, Shanghai Jiao Tong University, majoring in Computer Science and Technology. I am fortunate to be advised by Academician Weinan E and Prof. Junchi Yan. Before the Ph.D. program, I received my bachelor’s degree from SJTU and was upgraded from the master’s program of the Department of Computer Science and Engineering.
My research focuses on broad machine learning methodologies, especially machine learning for combinatorial optimization and decision making, generative models, and reasoning-oriented large language models. I have published 12 first-author/co-first-author papers at CCF-A top-tier conferences, including 10 papers at NeurIPS, ICML, and ICLR, with Spotlight recognitions at NeurIPS and ICML.
I have led multiple open-source projects on machine learning for discrete optimization, contributed to Huawei’s OptVerse-related technology initiatives, and participated in Alibaba’s RL4LLM framework development. My open-source projects have received over 5,000 stars in total, and the toolkits I led have accumulated over 70k downloads. I also serve as a reviewer for top-tier ML conferences (NeurIPS, ICML, ICLR, etc.) and journals (TPAMI, etc.).
Selected Experiences
-
T-Star Talent Program Intern, Alibaba ATH Business Group (June 2025 - Mar. 2026) Studied reinforcement-learning-based post-training for reasoning LLMs. Attention Illuminates LLM Reasoning (ICML 2026, first author) uses attention-revealed reasoning paths for targeted reward assignment and reached No. 1 on the Hugging Face daily paper leaderboard; Reasoning Palette (CVPR 2026, co-first author) enables controllable reasoning through latent contextualization; and FlowTracer (ICML 2026, co-first author) extracts an attention-induced information-flow backbone for token-level credit assignment. Contributed to Alibaba’s ALE agent learning ecosystem, the ROME agent model, and ROLL Flash: ROME matches 480B+ models in agentic coding with 3B activated parameters (30B total), while ROLL Flash achieves up to 2.72x speedup on agentic tasks and has received over 3k GitHub stars.
-
Researcher, ReThinkLab, Shanghai Jiao Tong University (July 2021 - Present) Developed generative machine learning paradigms and models, including PCL (ICML 2025) for label-repurposed supervised learning and AdvLatGAN (NeurIPS 2022 Spotlight) for latent-space optimization. Built generative combinatorial optimization frameworks including T2T (NeurIPS 2023), FastT2T (NeurIPS 2024), GenSCO (NeurIPS 2025), and MaskCO (ICLR 2026) as the first/co-first author, achieving substantial gains in solution quality and speed across TSP, MIS, Max Cut, etc. Proposed Unify ML4TSP (ICLR 2025) and led
awesome-ml4co,ML4CO-Kit,ML4TSPBench,ML4CO-Bench-101with over 3k GitHub stars and 70k downloads. -
Researcher, Learning-based Optimization Solver Project with Huawei Noah’s Ark Lab (Apr. 2022 - Feb. 2023) Developed generative data methods for SAT and MILP, including HardSATGEN (KDD 2023) and MixSATGEN (ICLR 2024). The methods were integrated into Huawei’s OptVerse AI Solver, which ranked first on the Hans Mittelmann benchmark and received the 2023 WAIC SAIL Award. Developed the Kissat-Adaptive-Restart solver for Huawei HiSilicon EDA scenarios, delivering an average performance improvement of about 18% and a maximum improvement of 97%.
Academic Performance
Undergraduate period:
- GPA: 91.03/100 (or 3.93/4.3), Rank: 3/129 (top 2.3%)
- Foundation Courses: 73.33% above A, 40.00% above A+
- Subject Courses: 80.00% above A, 50.00% above A+
Postgraduate period:
- GPA: 3.83/4.0
- Rank Reference: 3 out of 211 achieving the Graduate National Scholarship
- Courses: 90.0% at A level
Selected Awards
- Young Scientists (Ph.D) Fund of the National Natural Science Foundation of China (the only recipient in the school, the only F06 (AI category) recipient in the university, ¥300,000)
- Young Talent (Ph.D) Development Program of the China Association for Science and Technology (the only recipient in the school, ¥40,000)
- SJTU Pacemaker to Merit Student Award (top 10 university-wide)
- Graduate National Scholarship (top 1% in CS Dept.)
- Undergraduate National Scholarship (top 0.2% in the nation)
- Outstanding Graduate of Shanghai (top 3%)
- NeurIPS 2024 Top Reviewer Award (top 10%)
- Huawei Fellowship (top 3%)
- HyperGryph Fellowship (top 3%)
- 1st-Class Academic Excellence Scholarship (top 1%)
- Merit Student of Shanghai Jiao Tong University
- 1st-Class Academic Scholarship for Graduate Students
- Special Prize for Social Practice of SJTU
- First Prize for Social Practice of SJTU
- Advanced Individuals in Social Practice of SJTU
News
| Jun 3, 2026 | Eight papers were accepted by ICML/CVPR/ICLR 2026. |
|---|---|
| Oct 1, 2025 | Four paper was accepted by NeurIPS 2025. |
| May 1, 2025 | Two paper was accepted by ICML 2025. |
| Jan 23, 2025 | Two paper was accepted by ICLR 2025. |
| Sep 26, 2024 | Two paper was accepted by NeurIPS 2024. |
Publications
-
ICMLHow Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMsIn International Conference on Machine Learning, 2026
-
ICMLAttention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy OptimizationIn International Conference on Machine Learning, 2026
-
CVPRReasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMsIn Conference on Computer Vision and Pattern Recognition, 2026
-
ICLRMaskCO: Masked Generation Drives Effective Representation Learning and Exploiting for Combinatorial OptimizationIn International Conference on Learning Representations, 2026
-
ICLRConRep4CO: Contrastive Representation Learning of Combinatorial Optimization Instances across TypesIn International Conference on Learning Representations, 2026
-
ICLRNative Adaptive Solution Expansion for Diffusion-based Combinatorial OptimizationIn International Conference on Learning Representations, 2026
-
arXivLet It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem2026Core contributor
-
arXiv
-
NeurIPSBridging Crypto with ML-based Solvers: the SAT Formulation and BenchmarksIn Advances in Neural Information Processing Systems, 2025
-
中国科学Learning to Solve Combinatorial Optimization under Positive Linear Constraints via Non-Autoregressive Neural NetworksSCIENTIA SINICA Informationis 2024, 2024
