My research interests lie in large language model reasoning and agentic reinforcement learning, with a particular focus on GUI agents and autonomous AI scientists. My long-term goal is to build agents that can reason, learn from environmental feedback, and reliably solve complex real-world and scientific problems.
Situo Zhang*, Yifan Zhang*, Zichen Zhu, Da Ma, Lei Pan, Danyang Zhang, Zihan Zhao, Lu Chen, Kai Yu
We equip multimodal LLMs with image cropping and code-based computation, then train tool use through agentic reinforcement learning on DuoChart. This improves fine-grained visual grounding and numerical reasoning across chart benchmarks.
Situo Zhang*, Hanqi Li*, Lu Chen, Zihan Zhao, Xuanze Lin, Zichen Zhu, Bo Chen, Xin Chen, Kai Yu
We train a chemical reasoning LLM with reinforcement learning and chemically verifiable rewards for retrosynthesis prediction. Its explicit reasoning improves both accuracy and interpretability, reaching 65.0% top-1 accuracy on USPTO-50K.
Danyang Zhang*, Situo Zhang*, Ziyue Yang, Zichen Zhu, Zihan Zhao, Ruisheng Cao, Lu Chen, Kai Yu
We provide dense, step-level progress rewards for training GUI agents instead of scoring only final outcomes. An LCS-based self-annotation method identifies key trajectory steps and supplies progress labels without costly manual annotation.
Situo Zhang, Yifan Zhang, Zichen Zhu, Hankun Wang, Da Ma, Danyang Zhang, Lu Chen, Kai Yu
We dynamically adjust speculative draft length through lightweight blockwise pre-verification, stopping low-quality drafts before target-model verification. Pacer consistently improves standard speculative decoding and achieves up to 2.66× speedup over autoregressive decoding.