About

I am Ganlin Yang, a Ph.D. candidate at the University of Science and Technology of China and a joint PHD student at Shanghai AI Laboratory. My primary research interest is embodied intelligence, especially at the intersection of embodied manipulation and embodied brain models. I aim to develop agents that can perceive, reason, and act coherently in long-horizon tasks, with strong generalization across environments, tasks, and embodiments.

My recent work focuses on how multimodal foundation models can support embodied decision-making through unified end-to-end frameworks. Representative projects include VLASER, Visual Embodied Brain, and EventVLA. These works emphasize spatial intelligence, world-model-guided reasoning, and memory-enhanced policy learning for robust long-horizon behavior.

I also work on multimodal understanding and generation. I have contributed to the InternVL research line, including InternVL3.5, InternVL-U and Intern-S1. This line of work explores model capability scaling, multimodal reasoning, and large-scale open data/model pipelines. Before these directions, I also worked on 3D reconstruction and neural rendering. This experience provides a useful foundation for visual representation learning in embodied systems.

Education

University of Science and Technology of China
P.H.D candidate of Electronic Engineering & Information Science
School of Information Science and Technology
Supervisor: Prof. Jifeng Dai, Prof. Wengang Zhou
Sep. 2024 - Present
University of Science and Technology of China
Master of Electronic Engineering & Information Science
School of Information Science and Technology
Supervisor: Prof. Dong Liu
Sep. 2022 - Jun. 2024
University of Science and Technology of China
Bachelor of Electronic Engineering & Information Science
School of Gifted Young
GPA: 3.89/4.3 (ranked 10% at School of Gifted Young)
Sep. 2018 - Jun. 2022

Internship

Microsoft Research Asia (MSRA)
Research Intern, Intelligent Multimedia Group
Research Topic: 3D reconstruction and Neural Rendering
Supervisor: Dr. Zhizheng Zhang
August 2021 - July 2022
Microsoft Research Asia (MSRA)
Research Intern, Multimedia Computing Group
Research Topic: 3D reconstruction and generation
Supervisor: Dr. Jingjing Fu
July 2023 - June 2024
OpenGVLab, Shanghai AI Laboratory
Research Intern, Large Language Model Center
Research Topic: Multimodal Large Language Model
Supervisor: Dr. Jifeng Dai, Dr. Wenhai Wang
June 2024 - Nov. 2025
OpenRobotLab, Shanghai AI Laboratory
Research Intern, Physical Intelligence Center
Research Topic: Embodied AI; Vision Language Action Model
Supervisor: Dr. Jiangmiao Pang, Dr. Tai Wang
Nov. 2025 - Present
Shanghai Jiao Tong University
Visiting Intern, ScaleLab
Research Topic: Embodied AI; World-action Model
Supervisor: Dr. Yao Mu
Jan. 2026 - Present

Research Experiences

Embodied Brain and Manipulation

paper teaser

VLASER: Vision-Language-Action Model with Synergistic Embodied Reasoning

ICLR 2026 | [Paper] [GitHub] [Project Page]

Ganlin Yang*, Tianyi Zhang*, Haoran Hao*, Weiyun Wang, Yibin Liu, ..., Wenhai Wang, Yao Mu, Zhi Hou

Summary: VLASER introduces synergistic embodied reasoning to tightly couple scene understanding, instruction grounding, and action prediction for robust long-horizon manipulation.

paper teaser

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Technical Report | [Paper] [GitHub]

Gen Luo*, Ganlin Yang*, Ziyang Gong*, Guanzhou Chen*, ..., Yu Qiao, Rongrong Ji, Xizhou Zhu

Summary: This work proposes a visual embodied brain paradigm that unifies perception, spatial reasoning, and control planning to improve generality across embodied tasks and environments.

paper teaser

ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments

Technical Report | [Paper] [GitHub] [Project Page]

Ziyang Gong, Zehang Luo, Anke Tang, Zhe Liu, Shi Fu, Zhi Hou, Ganlin Yang, ..., Hengshuang Zhao, Dacheng Tao, Xiaogang Wang

Summary: ACE-Brain-0 argues for spatial intelligence as a common abstraction across embodiments, enabling transfer of planning and control priors between heterogeneous robots.

paper teaser

EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

Technical Report | [Paper] [GitHub] [Project Page]

Ganlin Yang*, Zhangzheng Tu*, Yuqiang Yang*, Sitong Mao, Junyi Dong, Tianxing Chen, Jiaqi Peng, Jing Xiong, Jiafei Cao, Jifeng Dai, Wengang Zhou, Yao Mu, Tai Wang.

Summary: EventVLA introduces event-driven memory updates to preserve key visual evidence during long-horizon interaction, improving temporal consistency and policy robustness.

paper teaser

Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

Technical Report | [Paper] [GitHub] [Project Page]

Jiaqi Peng*, Xiqian Yu*, Delin Feng*, ..., Ganlin Yang, ..., Jiangmiao Pang, Yuan Shen, Tai Wang.

Summary: Cortex is a bidirectionally aligned embodied agent framework with a customized planning interface that conveys executable and tractable subtask plans from high-level VLM to low-level VLA.

Multimodal Understanding and Generation

paper teaser

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Technical Report | [Paper] [GitHub] [Project Page]

Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, ... , Ganlin Yang, ... , Kai Chen, Yu Qiao, Wenhai Wang, Gen Luo.

Summary: InternVL3.5 systematically improves a unified multimodal model across perception, reasoning, and generation, with better scaling behavior and stronger efficiency-quality trade-offs.

paper teaser

InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing

Technical Report | [Paper] [GitHub]

Changyao Tian, Danni Yang, Guanzhou Chen, ... , Ganlin Yang, ... , Yu Qiao, Kai Chen, Hongjie Zhang.

Summary: InternVL-U presents a unified multimodal framework that supports both discriminative and generative tasks in one system, enabling broad capability transfer across modalities.

paper teaser

Intern-S1: A Scientific Multimodal Foundation Model

Technical Report | [Paper] [GitHub]

OpenGVLab Team.

Summary: Intern-S1 focuses on scientific multimodal understanding with stronger domain-oriented reasoning, aiming to bridge general MLLMs and scientific data-intensive applications.

paper teaser

ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework

Technical Report | [Paper] [GitHub]

Guanzhou Chen, Erfei Cui, Changyao Tian, Danni Yang, Ganlin Yang, Yu Qiao, Hongsheng Li, Gen Luo, Hongjie Zhang.

Summary: ScaleEdit-12M builds a multi-agent data engine to generate large-scale editing instruction data, improving data diversity and controllability for open-source image editing models.

3D Reconstruction & Rendering

paper teaser

Drim-NeRF: Diffusion-Based Restoration for Improving Neural Radiance Fields

TCSVT 2025 | [Paper]

Ganlin Yang, Kaidong Zhang, Jingjing Fu, Dong Liu.

Summary: Drim-NeRF introduces a diffusion-based restoration stage to refine degraded views and improve NeRF reconstruction quality under noisy, low-light, or sparsely sampled conditions.

paper teaser

Mask-Based Modeling for Neural Radiance Fields

ICLR 2024 | [Paper] [GitHub]

Ganlin Yang, Guoqiang Wei, Zhizheng Zhang, Yan Lu, Dong Liu.

Summary: This work proposes mask-guided modeling for NeRF training, improving geometry and appearance learning by focusing optimization on informative regions and reducing background-induced artifacts.

Skills

Programming:

Python, MATLAB, C/C++, PyTorch, LaTeX, Linux, Git, Deepspeed, Distributed Training, High-speed Model Inference

Research Skills:

3D Reconstruction and Perception, Multimodal Large Models Understanding and Generation, Reinforcement Learning for Embodied Control, Vision-Language-Action (VLA), World Action Model (WAM), Real-robot Deployment and Evaluation, End-to-end Embodied AI System Integration