About Me
I’m a Ph.D. student in the Department of Computer Science at University College London (UCL), advised by Professor Jun Wang. After years in industry, I have returned to academia to pursue my passion for research. I previously received my M.Sc. in Computational Statistics and Machine Learning (CSML) from UCL.
My research interests lie in Reinforcement Learning, Multi-Agent Systems, and Large Language Models.
Research Themes
How should an LLM-based agent decide what to learn from? My research explores three complementary notions of value for post-training agents, each suited to different forms of feedback and learning problems.
Verifiable Reward
Did it work?When outcomes can be checked, scalar rewards provide a dependable training signal for reasoning, decision-making, and multi-agent coordination.
Language Value
Why did it work?A scalar says how good an experience was; a language value function can also explain why, preserving structured knowledge that agents can reuse and refine.
Selected work Natural Language RL · Stateful Predictive KnowledgeDiscovery Value
What should we try next?For open-ended problems, discovery value can encode whichever signals matter for choosing what to generate, refine, or test next, such as performance, uncertainty, novelty, or information gain. In Large Discovery Models, a Gaussian-process acquisition function is one concrete instantiation.
Selected work Large Discovery Models
If you’d like to discuss potential collaborations or shared research interests, feel free to contact me at yan.song.24[at]ucl.ac.uk.
News
[2026.09] Our paper ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents has been accepted by EMNLP 2026! [Paper]
[2026.08] Our new paper Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search is now available on arXiv! It targets the problem that language models don’t know how to explore or exploit, an everlasting topic for RL scienctists. [Paper]
[2026.07] Our paper Learning Stateful Predictive Knowledge From Experience is available at the ICML 2026 AIWILD Workshop. [Paper] [Blog] [X Post]
[2026.05] Here comes our second collaboration paper with Li Auto: The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design and has been accepted by ICML 2026 !
[2026.02] We have been closely collaborating with Li Auto on several research topics. Now we have released our first joint paper: Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs. Well Done Guys! Stay tuned for more to come out!
[2025.05] Our paper Ask more, know better: Reinforce-Learned Prompt Questions for Decision Making with Large Language Models got accepted by ECML-PKDD 2025. A Testament to Persistence!
[2025.05] We have successfully held the AAMAS 2025 Online AI Competitions!
[2025.03] RL can now interactively train two LLM agents to reason ! Our paper ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning is available on Arxiv ! (Neurips 2025)
[2025.01] Our paper Efficient Reinforcement Learning with Large Language Model Priors got accepted by ICLR 2025 !
[2024.10] We release our LLM reasoning framework – OpenR !
