About me
Hi, I am Wish (Peter) Suharitdamrong.
I’m a PhD student at the Surrey Institute for People-Centred AI (PAI), University of Surrey, supervised by Dr. Sara Atito, Dr. Muhammad Awais and Dr. Xiatian Zhu. My research focuses on multimodal foundation models across vision, audio, and language, particularly multimodal large language models and multimodal representation learning. I previously completed an MSc in Artificial Intelligence at Surrey, where my master’s project was also supervised by Dr. Muhammad Awais and co-supervised by Tony Alex. Before that I completed a BSc in Computer Science at Surrey, supervised by Prof. Zhenhua Feng.
I’ve been working on multimodal problems since the start of my deep learning journey: my undergraduate dissertation focused on audio-visual talking face generation, and my master’s dissertation explored parameter-efficient fine-tuning of vision-language foundation models for multi-task visual grounding.
Research interests
- Multimodal Large Language Models (MLLMs)
- Representation Learning
News
- 2026-04 🥤 CoLA accepted at ICML 2026 🔥
Selected work
- 👀 The Hidden Evolution of Disguised Visual Context inside the VLM A study of how different vision integration paradigms behave in VLMs, and how each one transforms the visual representation inside the LLM.arXiv preprint (under review), 2026
- 🥤 CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks A PEFT framework that extends LoRA with a dedicated inter-modal adaptation pathway, for adapting dual-stream unimodal encoders to multimodal tasks.ICML 2026
Service
- Reviewer AAAI 2025 · AAAI 2026 · EMNLP 2026
Elsewhere
- GitHub @peterwisu
- X / Twitter @peterwisu
- LinkedIn wish-suharitdamrong
Contact
- personal peterwisu@gmail.com
- university ws00372@surrey.ac.uk