Zhongquan Zhou
Computer Science @ Shanghai Jiao Tong University
I am a master's student majoring in Computer Science at SJTU. My current interests include Agentic RL, AI Infrastructure, and Generative Algorithms.
Education
2025.09 – 2028.03
Shanghai Jiao Tong University
M.S. student, Computer Science. GPA: 3.70/4.0.
2021.09 – 2025.06
Lanzhou University
B.S., Computer Science. GPA: 3.97/5.0.
Research Experience
2026
CM-OPD: Conflict-Aware Multi-Teacher On-Policy Distillation
- Scope and ownership: Transferred mathematical, coding, and medical capabilities from three specialist teachers into one student model, leading teacher fine-tuning, on-policy rollout, token scoring, conflict-aware aggregation, and reproducibility tooling.
- Conflict-aware distillation: Had the student sample its own completions and frozen teachers score the same token sequences, then aggregated teacher–student log-probability differences with equal-weight and conflict-aware variants.
- Evaluation and findings: Used fixed-protocol GSM8K, MBPP, MedQA-zh, C-Eval, and IFEval benchmarks. Full OPD achieved the highest macro average (68.27), while Equal OPD remained stronger on GSM8K and MBPP.
2026.02 – 2026.06
DeepMatter: Autoregressive Generative Foundation Model for Crystal Materials
- Project focus: Conducted crystal-domain pretraining, property control, and GRPO reinforcement learning; the paper is under review at Nature Machine Intelligence.
- Data and pretraining: Built approximately 2 million crystal samples and trained a 100M-parameter autoregressive base model for unified property prediction and controllable generation.
- Property control and GRPO: Introduced continuous property tokens for stability and band-gap conditioning, then trained a property reward model to optimize structure–property consistency with GRPO.
- Diversity and results: Used an element-distribution buffer and diversity penalties to mitigate mode collapse; achieved stronger property prediction than traditional GNNs and improved unconditional/conditional generation over MatterGen.
2026
Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction
Mingze Kong, Zikun Qu, Zhongquan Zhou, et al. ICLR 2026 Workshop Accept.
- Problem formulation: Studied multi-turn workflow construction as policy optimization over grouped action subsequences.
- Optimization: Explored group sub-sequence policy optimization to improve credit assignment across intermediate workflow steps.
- Evaluation: Analyzed workflow generation quality with an emphasis on coherent and reliable multi-turn execution.
2025
ACORN: Adaptive Contrastive Optimization for Safe and Robust Fine-Grained Robotic Manipulation
Zhongquan Zhou, Shuhao Li, Zixian Yue.
- Method: Proposed ACORN, an adaptive contrastive optimization algorithm for robust and safe robotic manipulation.
- Training: Used dual-perturbation contrastive learning to align robot trajectories with expert demonstrations without manual annotation.
- Evaluation: Designed safety-centered metrics to quantify manipulation stability under extreme conditions.
Projects
2025.09 – 2025.12Shanghai AI Lab
Equivariant Machine Learning Force Fields and Generative Materials Design
- Generative modeling: Studied diffusion models and flow matching for materials generation with energy-guided vector field optimization.
- Equivariant modeling: Explored three-body attention to strengthen representations of complex atomic systems.
- Efficiency: Investigated low-rank decomposition and related acceleration techniques for large-scale materials pretraining.
2024.08 – 2025.03
Few-Shot AI System for Industrial Automation Control with Transfer Learning and Physics-Based Interpretability
- Time-aware control: Modeled time-dependent industrial processes and generated full-cycle schedules for temperature, aeration, and material ratios.
- Few-shot closed loop: Combined transfer learning, process knowledge, and sensor feedback for explainable real-time prediction and autonomous control.
- Deployment: Integrated hundreds of process sensors with a lightweight, domestic-GPU-compatible stack to improve stability and reduce data and compute costs.
Honors & Awards
- 2025.05 Outstanding Graduate of Gansu Province
- 2025.05 Honorable Mention, Mathematical Contest in Modeling 2025
- 2024.12 National Scholarship for Undergraduates, 2023-2024
- 2024.12 Outstanding Graduate of Lanzhou University
- 2024.07 National Third Prize, 15th China Students Service Outsourcing Innovation and Entrepreneurship Competition
- 2023.12 Lanzhou University Xiaomi First-Class Scholarship
- 2023.11 Lanzhou University-Huawei "Intelligent Foundation" Scholarship
- 2023.10 Second Prize, Gansu Division, CUMCM 2023
- 2023.07 National Second Prize, 16th Chinese Collegiate Computing Competition
- 2023.06 Outstanding Assistant Class Advisor of Lanzhou University
- 2023.05 Outstanding Student Cadre of Lanzhou University
- 2022.12 First-Class Scholarship for Outstanding Students, Lanzhou University
- 2022.10 National First Prize, Code Review Track, 5th China Open Source Innovation Competition