Ruoteng Li

李若腾

I am a Senior Applied Scientist at Microsoft.

Previously, I was an Applied Scientist at Amazon and a Research Scientist at ByteDance Singapore, where I worked on video understanding, multi-modal learning, and video recommendation. I was also a joint Lecturer at NUS from 2022 to 2023.

I obtained my PhD degree from the Electrical & Computer Engineering department at NUS in 2020 (2016-2020), supervised by Prof. Cheong Loong Fah and Prof. Robby T. Tan, who gave me invaluable guidance and indispensable encouragement. I was a research assistant at Yale-NUS (2016-2019) and a visiting researcher at the Cambridge Image Analysis Group, University of Cambridge. I have completed internships at Google DayDream ARCore teams and Tencent ARC Lab.

Before that, I completed my bachelor's degree from NUS and my high school education at Yao Hua High School in China.

NUS Yale-NUS Cambridge ByteDance Google TikTok Amazon
Ruoteng Li

Research Interest

Image/Video Generation, Diffusion Models, Vision-Language Models, Video Understanding, Multi-Modal Learning, and Image processing.

Academic Service

Area Chair: NeurIPS 2026.

Reviewer: CVPR (2022, 2023, 2024, 2026), ECCV (2024), ICCV (2023), ICLR (2022, 2023, 2024, 2025, 2026), NeurIPS (2022, 2023, 2024, 2025), AAAI (2022, 2023, 2024, 2026), WACV (2022, 2024, 2025, 2026), ACCV (2024), IJCV (2020, 2021, 2022, 2023).

Publications

Method schematics summarize each paper. Select a schematic to enlarge it.

SVP: generate captions, use grounding feedback to refine them, select informative text and fine-tune the VLM.
Feedback-Driven Vision-Language Alignment via Sampling-based Visual Projection
Giorgio Giannone, Evgeny Perevodchikov, Qianli Feng, Ruoteng Li, Rui Chen, Aleix Martinez
Transactions on Machine Learning Research, TMLR, 2026
TMLR 2026
Novel View Synthesis
Detail Preserving Novel View Synthesis
Yuelong Li, Ruoteng Li, et al.
In Submission, 2025
In Submission
LSFD: stereo movie frames produce refined depth for realistic paired hazy and clean training images.
Realistic Large-Scale Fine-Depth Dehazing Dataset from 3D Videos
Ruoteng Li, Xiaoyi Zhang, Shaodi You, Yu Li
arXiv preprint
arXiv
Reflection removal: multi-image cross-reconstruction trains a shared network to separate layers from one image.
Single Image Reflection removal via learning with multi-image constraints
Yingda Yin, Qingnan Fan, Dongdong Chen, Yujie Wang, Angelica Aviles-Rivero, Ruoteng Li, Carola-Bibiane Schönlieb, Baoquan Chen
arXiv preprint
arXiv
CREPE: alternate graph 1-Laplacian pseudo-label inference with deep classifier training.
Energy Models for Better Pseudo-Labels: Improving Semi-Supervised Classification with the 1-Laplacian Graph Energy
Angelica I. Aviles-Rivero, Nicolas Papadakis, Ruoteng Li, Philip Sellars, Samar M Alsaleh, Robby T Tan, Carola-Bibiane Schönlieb
In Submission to TPAMI
arXiv
Reflectance guidance and shadow/specular attention recover diffuse material color.
Estimating reflectance layer from a single image: Integrating reflectance guidance and shadow/specular aware learning
Yeying Jin, Ruoteng Li, Wenhan Yang, Robby T. Tan
AAAI Conference on Artificial Intelligence, AAAI, 2023
AAAI 2023
Hybrid warping: fuse forward and backward warped frames, features and edges to synthesize an intermediate frame.
Hybrid Warping Fusion for Video Frame Interpolation
Yu Li, Ye Zhu, Ruoteng Li, Xintao Wang, Yue Luo, Ying Shan
International Journal of Computer Vision, IJCV, 2022
IJCV 2022
Camera-compensated trajectory prediction helps track an object through occlusion.
Object Tracking using Spatio-Temporal Networks for Future Prediction Location
Ruoteng Li*, Yuan Liu*, Yu Cheng, Robby T. Tan, Xiubao Sui (* equal contribution)
European Conference on Computer Vision, ECCV, 2020
ECCV 2020
Weather-specific encoders, architecture search and a shared decoder restore images across weather types.
All-in-One Bad Weather Removal using Fusion Search
Ruoteng Li, Robby T. Tan, Loong-Fah Cheong
Conference on Computer Vision and Pattern Recognition, CVPR, 2020
CVPR 2020
RainFlow: learned veiling-invariant and streak-invariant features support optical flow estimation.
RainFlow: Optical Flow under Rain Streaks and Rain Veiling Effect
Ruoteng Li, Robby T. Tan, Loong-Fah Cheong, Angelica I. Aviles-Rivero, Qingnan Fan, Carola-bibiane Schönlieb
International Conference on Computer Vision, ICCV, 2019
ICCV 2019
GraphXNET: propagate a few chest X-ray labels across a graph of images using 1-Laplacian energy.
GraphXNET - Chest X-Ray Classification Under Extreme Minimal Supervision
Angelica I. Aviles-Rivero, Nicolas Papadakis, Ruoteng Li, Philip Sellars, Qingnan Fan, Robby T. Tan, Carola-bibiane Schönlieb
Medical Image Computing and Computer Assisted Intervention, MICCAI, 2019
MICCAI 2019
Estimate streaks, transmission and atmospheric light, then refine the reconstruction with a depth-guided GAN.
Heavy Rain Image Restoration: Integrating Physics Model and Conditional Adversarial Learning [code]
Ruoteng Li, Loong-Fah Cheong, Robby T. Tan
Conference on Computer Vision and Pattern Recognition, CVPR, 2019
CVPR 2019
Combine colored residue images and smooth structure layers for robust optical flow in rain.
Robust Optical Flow in Rainy Scenes [code]
Ruoteng Li, Robby T. Tan, Loong-Fah Cheong
European Conference on Computer Vision, ECCV, 2018
ECCV 2018
SMRNet: scale-specific recurrent branches and a veil module restore rainy images in multiple stages.
Single Image Deraining using Scale-Aware Multi-Stage Recurrent Network
Ruoteng Li, Loong-Fah Cheong, Robby T. Tan
arXiv preprint, 2017
arXiv

Teaching

NUS

National University of Singapore

EE4704 - Image Processing and Analysis

2022 - 2023