I am a postdoctoral researcher at TU Delft, working with Frans A. Oliehoek and Jan-Willem van de Meent @UvA. I am also a visiting researcher at MPI-IS, supervised by Shiwei Liu. I obtained my Ph.D. from TU/e, supervised by Mykola Pechenizkiy and Meng Fang, and I am very fortunate to have worked closely with Yali Du @KCL and Biwei Huang @UCSD. Previously, I was a research intern at Microsoft, and I obtained my Master’s and Bachelor’s degrees at SDU.
My research focuses on learning agents: agents that improve from experience and feedback.
- LLM agents: self-evolving agents, agents with symbolic / logic rule knowledge
- Classic RL: learning from episodic rewards, causal reward redistribution
Open to new positions, visiting opportunities, and collaborations.
🎓 Education
Sep 2022 – Sep 2026
Ph.D. in Computer Science, TU/e · Supervisors: Mykola Pechenizkiy, Meng Fang
Sep 2019 – Jul 2022
M.Sc. in Control Science and Engineering, SDU · Supervisor: Wei Zhang
Sep 2015 – Jul 2019
B.Sc. in Automation, SDU
🧑💻 Experience
Sep 2026 – present
TU Delft, Postdoctoral researcher
Supervisors: Frans A. Oliehoek, Jan-Willem van de Meent
Supervisors: Frans A. Oliehoek, Jan-Willem van de Meent
Apr 2026 – present
MPI-IS, Visiting researcher · Supervisor: Shiwei Liu
Mar – Oct 2024
Microsoft, Research intern · Mentor: Lu Wang
✨ News
Sep 2026
Joined Delft University of Technology as a postdoctoral researcher.
May 2026
One paper was accepted by RLC 2026. See you in Canada!
May 2026
One paper was accepted by ICML 2026. See you in Korea!
Apr 2026
Started a new journey at the Max Planck Institute for Intelligent Systems.
Feb 2026
One paper was accepted by AAMAS 2026.
📄 Publications Full list at Google Scholar · * co-first author
Preprints
- A Causal Approach for Interpretable Reward Redistribution in Visual Reinforcement Learning
Under Review - Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?
Preprint Link
Conferences
- CAST: A Causality-inspired Spatial-temporal Return Decomposition Approach for Multi-agent RL.
RLC 2026 Link - Self-Evolving LLM Agents with In-distribution Optimization.
ICML 2026 Link - Causality Meets Locality: Provably Generalizable and Scalable Policy Learning for Networked Systems.
NeurIPS 2025 Link - Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-level Physics Problem Solving.
EMNLP Findings 2025 Link - RuAG: Learned-rule-augmented Generation for Large Language Models.
ICLR 2025 Link - Pillagerbench: Benchmarking LLM-based Agents in Competitive Minecraft Team Environments.
IEEE Conference on Games (CoG), 2025 Link - Large Language Models are Neurosymbolic Reasoners.
AAAI 2024 Link - Interpretable Reward Redistribution in Reinforcement Learning: A Causal Approach.
NeurIPS 2023 Link - COOM: A Game Benchmark for Continual Reinforcement Learning.
NeurIPS 2023 D&B Track Link - RSPT: Reconstruct Surroundings and Predict Trajectories for Generalizable Active Object Tracking.
AAAI 2023 Link
