About Me

I’m a Ph.D. student in the Department of Computer Science at University College London (UCL), advised by Professor Jun Wang. After years in industry, I have returned to academia to pursue my passion for research. I previously received my M.Sc. in Computational Statistics and Machine Learning (CSML) from UCL.

My research interests lie in Reinforcement Learning, Multi-Agent Systems, and Large Language Models.

Research Themes

How should an LLM-based agent decide what to learn from? My research explores three complementary notions of value for post-training agents, each suited to different forms of feedback and learning problems.

  • Verifiable Reward

    Did it work?

    When outcomes can be checked, scalar rewards provide a dependable training signal for reasoning, decision-making, and multi-agent coordination.

  • Language Value

    Why did it work?

    A scalar says how good an experience was; a language value function can also explain why, preserving structured knowledge that agents can reuse and refine.

  • Discovery Value

    What should we try next?

    For open-ended problems, discovery value can encode whichever signals matter for choosing what to generate, refine, or test next, such as performance, uncertainty, novelty, or information gain. In Large Discovery Models, a Gaussian-process acquisition function is one concrete instantiation.

    Selected work Large Discovery Models

If you’d like to discuss potential collaborations or shared research interests, feel free to contact me at yan.song.24[at]ucl.ac.uk.


News