RESEARCHER · JD FUTURE ACADEMY
I work on efficient large language models, diffusion-based generative systems, and unified multimodal intelligence — from the optimization underneath training to the architectures that shape images, video, and reasoning.
Previously, I received my M.S. from Peking University under the supervision of Prof. Zaiwen Wen, and my B.S. from Jiangnan University.
01Large Language Models02Diffusion Models03Multimodal Systems
CURRENT FOCUS
LLM × DIFFUSION × MULTIMODALMaking generative models more efficient, scalable, and controllable.
01 — PUBLICATIONS
Selected publications
Research on efficient optimization, language-model reasoning, and continuous generative modeling.
PEER-REVIEWED & SELECTED WORK
Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning
Generative Pre-trained Autoregressive Diffusion Transformer
A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models
Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures
TECHNICAL REPORTCORE CONTRIBUTOR
* Equal contribution
02 — BLOG
Notes & ideas
A space for research notes, technical walkthroughs, and ideas that are still taking shape.
01
Writing in progress.
The first post will appear here.
03 — JOURNEY