About Me

Zhongquan Zhou

Zhongquan Zhou

Computer Science @ Shanghai Jiao Tong University

I am a master's student majoring in Computer Science at SJTU. My current interests include Agentic RL, AI Infrastructure, and Generative Algorithms.

Email: zhouzhq2021@sjtu.edu.cn GitHub: zhouzhq2021 Location: Shanghai, China

Education

2025.09 – 2028.03

Shanghai Jiao Tong University

M.S. student, Computer Science. GPA: 3.70/4.0.

2021.09 – 2025.06

Lanzhou University

B.S., Computer Science. GPA: 3.97/5.0.

Research Experience

2026

CM-OPD: Conflict-Aware Multi-Teacher On-Policy Distillation

  • Scope and ownership: Transferred mathematical, coding, and medical capabilities from three specialist teachers into one student model, leading teacher fine-tuning, on-policy rollout, token scoring, conflict-aware aggregation, and reproducibility tooling.
  • Conflict-aware distillation: Had the student sample its own completions and frozen teachers score the same token sequences, then aggregated teacher–student log-probability differences with equal-weight and conflict-aware variants.
  • Evaluation and findings: Used fixed-protocol GSM8K, MBPP, MedQA-zh, C-Eval, and IFEval benchmarks. Full OPD achieved the highest macro average (68.27), while Equal OPD remained stronger on GSM8K and MBPP.
2026.02 – 2026.06

DeepMatter: Autoregressive Generative Foundation Model for Crystal Materials

  • Project focus: Conducted crystal-domain pretraining, property control, and GRPO reinforcement learning; the paper is under review at Nature Machine Intelligence.
  • Data and pretraining: Built approximately 2 million crystal samples and trained a 100M-parameter autoregressive base model for unified property prediction and controllable generation.
  • Property control and GRPO: Introduced continuous property tokens for stability and band-gap conditioning, then trained a property reward model to optimize structure–property consistency with GRPO.
  • Diversity and results: Used an element-distribution buffer and diversity penalties to mitigate mode collapse; achieved stronger property prediction than traditional GNNs and improved unconditional/conditional generation over MatterGen.
2026

Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction

Mingze Kong, Zikun Qu, Zhongquan Zhou, et al. ICLR 2026 Workshop Accept.

  • Problem formulation: Studied multi-turn workflow construction as policy optimization over grouped action subsequences.
  • Optimization: Explored group sub-sequence policy optimization to improve credit assignment across intermediate workflow steps.
  • Evaluation: Analyzed workflow generation quality with an emphasis on coherent and reliable multi-turn execution.
2025

ACORN: Adaptive Contrastive Optimization for Safe and Robust Fine-Grained Robotic Manipulation

Zhongquan Zhou, Shuhao Li, Zixian Yue.

  • Method: Proposed ACORN, an adaptive contrastive optimization algorithm for robust and safe robotic manipulation.
  • Training: Used dual-perturbation contrastive learning to align robot trajectories with expert demonstrations without manual annotation.
  • Evaluation: Designed safety-centered metrics to quantify manipulation stability under extreme conditions.

Projects

2025.09 – 2025.12Shanghai AI Lab

Equivariant Machine Learning Force Fields and Generative Materials Design

  • Generative modeling: Studied diffusion models and flow matching for materials generation with energy-guided vector field optimization.
  • Equivariant modeling: Explored three-body attention to strengthen representations of complex atomic systems.
  • Efficiency: Investigated low-rank decomposition and related acceleration techniques for large-scale materials pretraining.
2024.08 – 2025.03

Few-Shot AI System for Industrial Automation Control with Transfer Learning and Physics-Based Interpretability

  • Time-aware control: Modeled time-dependent industrial processes and generated full-cycle schedules for temperature, aeration, and material ratios.
  • Few-shot closed loop: Combined transfer learning, process knowledge, and sensor feedback for explainable real-time prediction and autonomous control.
  • Deployment: Integrated hundreds of process sensors with a lightweight, domestic-GPU-compatible stack to improve stability and reduce data and compute costs.

Honors & Awards

  • 2025.05 Outstanding Graduate of Gansu Province
  • 2025.05 Honorable Mention, Mathematical Contest in Modeling 2025
  • 2024.12 National Scholarship for Undergraduates, 2023-2024
  • 2024.12 Outstanding Graduate of Lanzhou University
  • 2024.07 National Third Prize, 15th China Students Service Outsourcing Innovation and Entrepreneurship Competition
  • 2023.12 Lanzhou University Xiaomi First-Class Scholarship
  • 2023.11 Lanzhou University-Huawei "Intelligent Foundation" Scholarship
  • 2023.10 Second Prize, Gansu Division, CUMCM 2023
  • 2023.07 National Second Prize, 16th Chinese Collegiate Computing Competition
  • 2023.06 Outstanding Assistant Class Advisor of Lanzhou University
  • 2023.05 Outstanding Student Cadre of Lanzhou University
  • 2022.12 First-Class Scholarship for Outstanding Students, Lanzhou University
  • 2022.10 National First Prize, Code Review Track, 5th China Open Source Innovation Competition

Skills

Python PyTorch Machine Learning LLM Post-train Diffusion Models Streamlit Linux CET-4: 565 CET-6: 520 Mandarin: Level 2-A