Zikang Shan(单子康)

I am a second-year Ph.D. student at Peking University, advised by Prof. Liwei Wang. I'm fortunate to also work closely with Prof. Han Zhong. Currently, I am a research intern at Qwen Business Unit; previously, I spent two years as a research intern at Microsoft Research Asia, mentored by Dr. Li Zhao. I received my Bachelor's degree from Peking University. During my undergraduate years, I was honored to be advised by Prof. Liwei Wang and Prof. He Wang.

Please feel free to contact me if you want to discuss or collaborate! Email: shanzikang [at] stu.pku.edu [dot] cn

Google Scholar  |  ORCID  |  Github  |  X  |  Linkedin

News

Aug, 2026 Joined Qwen Business Unit as an LLM post-training algorithm intern.
Apr, 2026 Paper available on Arxiv.
Apr, 2026 Paper accepted at ACL main conference.
May, 2025 Paper accepted at ICML as spotlight. See you in Vancouver!
Sep, 2024 Joined Microsoft Research Asia as a research intern.

Research

My research focuses on the credit assignment problem in reinforcement learning for LLM post-training: how to attribute sparse outcome feedback to token-level decisions. I consider a principled, general and scalable solution as the key to robust scaling of LLM RL, and have pursued it through different approaches: deriving implicit process reward models, designing model-free hierarchical advantage estimation, and improving value model expressivity. As this bottleneck grows more pressing in frontier agentic RL training, I am actively pursuing better solutions.

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
Zikang Shan, Han Zhong, Liwei Wang, Li Zhao
Under review
Paper
Grounded by representation complexity and approximation experiments, we identify limited critic expressivity as a reason why PPO fails in reasoning tasks, and propose generative critics that perform chain-of-thought reasoning before prediction, improving both value estimation and downstream RL performance.
SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning
Zhengyang Ai, Zikang Shan, Xiaodong Ai, Jingxian Tang, Hangkai Hu, Pinyan Lu
ACL 2026, main conference
Paper
We propose a model-free fine-grained credit assignment framework that rewards chain-of-thought segments with a tractable proxy and redistributes intra-segment credit by token entropy, yielding increased performance and token efficiency.
DPO Meets PPO: Reinforced Token Optimization for RLHF
Han Zhong*, Zikang Shan*, Guhao Feng*, Wei Xiong, Xinle Cheng, Li Zhao, Di He, Jiang Bian, Liwei Wang (* equal contribution)
ICML 2025, spotlight poster
Paper  |  Code
We derive DPO-aligned models as implicit token-level process reward models and optimize with PPO, achieving theoretically grounded fine-grained alignment with improved sample efficiency and final performance.
UniDexGrasp++: Improving Universal Dexterous Grasping via Geometry-aware Curriculum Learning and Iterative Generalist-Specialist Learning
Weikang Wan*, Haoran Geng*, Yun Liu, Zikang Shan, Li Yi, Yaodong Yang, and He Wang
ICCV 2023, oral presentation, all top rankings, best paper finalist
Paper  |  Website  |  Code
We improve our previous method, making it object-agnostic and much more effective.
UniDexGrasp: Universal Robotic Dexterous Grasping via Learning Diverse Proposal Generation and Goal-Conditioned Policy
Yinzhen Xu*, Weikang Wan*, Jialiang Zhang*, Haoran Liu*, Zikang Shan, Hao Shen, Ruicheng Wang, Haoran Geng, Yijia Weng, Jiayi Chen, Tengyu Liu, Li Yi, and He Wang
CVPR 2023
Paper  |  Website  |  Code
We propose a method to learn dexterous grasping policies able to handles diverse objects based on realistic observations.

This website is adapted from Jon Barron's website and deployed on Github Pages. Last updated on Aug 2025.