I am a postdoctoral researcher at TU Delft, working with Frans A. Oliehoek and Jan-Willem van de Meent @UvA. I am also a visiting researcher at MPI-IS, supervised by Shiwei Liu. I obtained my Ph.D. from TU/e, supervised by Mykola Pechenizkiy and Meng Fang, and I am very fortunate to have worked closely with Yali Du @KCL and Biwei Huang @UCSD. Previously, I was a research intern at Microsoft, and I obtained my Master’s and Bachelor’s degrees at SDU.

My research focuses on learning agents: agents that improve from experience and feedback.

  • LLM agents: self-evolving agents, agents with symbolic / logic rule knowledge
  • Classic RL: learning from episodic rewards, causal reward redistribution

Open to new positions, visiting opportunities, and collaborations.

🎓 Education
Sep 2022 – Sep 2026
Ph.D. in Computer Science, TU/e · Supervisors: Mykola Pechenizkiy, Meng Fang
Sep 2019 – Jul 2022
M.Sc. in Control Science and Engineering, SDU · Supervisor: Wei Zhang
Sep 2015 – Jul 2019
B.Sc. in Automation, SDU
🧑‍💻 Experience
Sep 2026 – present
TU Delft, Postdoctoral researcher
Supervisors: Frans A. Oliehoek, Jan-Willem van de Meent
Apr 2026 – present
MPI-IS, Visiting researcher · Supervisor: Shiwei Liu
Mar – Oct 2024
Microsoft, Research intern · Mentor: Lu Wang
✨ News
Sep 2026
Joined Delft University of Technology as a postdoctoral researcher.
May 2026
One paper was accepted by RLC 2026. See you in Canada!
May 2026
One paper was accepted by ICML 2026. See you in Korea!
Apr 2026
Started a new journey at the Max Planck Institute for Intelligent Systems.
Feb 2026
One paper was accepted by AAMAS 2026.

All news →

📄 Publications Full list at Google Scholar  ·  * co-first author

Preprints

  • A Causal Approach for Interpretable Reward Redistribution in Visual Reinforcement Learning Causal RL Under Review
    Yudi Zhang, Yali Du, Biwei Huang, Ziyan Wang, Jun Wang, Meng Fang, Mykola Pechenizkiy. Under Review
  • Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones? LLM Agents Preprint
    Yudi Zhang, Lu Wang, Meng Fang, Yali Du, et al. Preprint Link

Conferences

  • CAST: A Causality-inspired Spatial-temporal Return Decomposition Approach for Multi-agent RL. Causal RL RLC 2026
    Yudi Zhang, Yali Du, Biwei Huang, Meng Fang, Mykola Pechenizkiy. RLC 2026 Link
  • Self-Evolving LLM Agents with In-distribution Optimization. LLM Agents ICML 2026
    Yudi Zhang, Meng Fang, Zhenfang Chen, Mykola Pechenizkiy. ICML 2026 Link
  • Causality Meets Locality: Provably Generalizable and Scalable Policy Learning for Networked Systems. Causal RL NeurIPS 25 Spotlight
    Hao Liang*, Shuqing Shi*, Yudi Zhang, Biwei Huang, Yali Du. NeurIPS 2025 Link
  • Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-level Physics Problem Solving. LLM RAG EMNLP Findings 25
    Shunfeng Zheng*, Yudi Zhang*, Meng Fang, Zihan Zhang, Zhitan Wu, Mykola Pechenizkiy, Ling Chen. EMNLP Findings 2025 Link
  • RuAG: Learned-rule-augmented Generation for Large Language Models. LLM Agents ICLR 25
    Yudi Zhang*, Pei Xiao*, Lu Wang, Chaoyun Zhang, Meng Fang, Yali Du, et al. ICLR 2025 Link
  • Pillagerbench: Benchmarking LLM-based Agents in Competitive Minecraft Team Environments. LLM Agents IEEE CoG 25 Oral
    Olivier Schipper, Yudi Zhang, Yali Du, Mykola Pechenizkiy, Meng Fang. IEEE Conference on Games (CoG), 2025 Link
  • Large Language Models are Neurosymbolic Reasoners. LLM Agents AAAI 24
    Meng Fang*, Shilong Deng*, Yudi Zhang*, Zijing Shi, Ling Chen, Mykola Pechenizkiy, Jun Wang. AAAI 2024 Link
  • Interpretable Reward Redistribution in Reinforcement Learning: A Causal Approach. Causal RL NeurIPS 23
    Yudi Zhang, Yali Du, Biwei Huang, Ziyan Wang, Jun Wang, Meng Fang, Mykola Pechenizkiy. NeurIPS 2023 Link
  • COOM: A Game Benchmark for Continual Reinforcement Learning. Continual RL NeurIPS D&B 23
    Tristan Tomilin, Meng Fang, Yudi Zhang, Mykola Pechenizkiy. NeurIPS 2023 D&B Track Link
  • RSPT: Reconstruct Surroundings and Predict Trajectories for Generalizable Active Object Tracking. Embodied AI AAAI 23 Oral
    Fangwei Zhong*, Xiao Bi*, Yudi Zhang, Wei Zhang, Yizhou Wang. AAAI 2023 Link

Journals

  • Large action models: From inception to implementation. LLM Agents TMLR
    TMLR 2025 Link
  • MACCA: Offline Multi-agent Reinforcement Learning with Causal Credit Assignment. Causal RL TMLR
    Ziyan Wang, Yali Du, Yudi Zhang, Meng Fang, Biwei Huang. TMLR 2025 Link