Hi there!


Geng Li (李耕)

avatar

I am a Ph.D. student in the 2023 cohort at the Wangxuan Institute of Computer Technology, Peking University, advised by Prof. Yuxin Peng. My research focuses on fine-grained perception and visual reasoning in multimodal large language models. As a first author, I have published several papers at top-tier international conferences, including CVPR, ICML, and ECCV. My representative works include DyFo, BVS, and DiCoBench. DyFo was selected as a CVPR 2025 Highlight paper, with a Highlight selection rate of only 13.5%. My work has received dedicated coverage from several academic platforms, including VALSE, CVer, and Jishi Community. I also received the VALSE 2025 Popular Poster Award, ranking among the top 11 out of 398 posters.

Latest Publications

  [ICML 2026BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception

  [ECCV 2026DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues

  [CVPR 2025 (Highlight)DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding

  [ICLR 2025Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models

  [TKDE 2023Automated Graph Neural Network Search Under Federated Learning Framework

  [WISE 2022EEML: Ensemble Embedded Meta-Learning