My name is Lin Zhu. I received a B.S. degree in Statistics from Central South University, China, in 2019 and an M.S. degree in Statistics from Shandong University, China, in 2022. I recently completed my Ph.D in Computer Science and Technology at Shanghai Jiao Tong University, China, advised by Prof.Nanyang Ye.

My research interests include Multimodal Perception and Reasoning, Trustworthy AI, and Out-of-Distribution Generalization. I have published several papers on top machine learning journals and conferences, including TPAMI, IJCV, NeurIPS, ICML, ICLR, and AAAI.

If you are interested in potential research collaborations, please feel free to contact me via email.

👐 Recent Updates

  • 2026.01: One journal paper CFSM accepted to TPAMI.
  • 2025.08: One journal paper InfoBound accepted to TPAMI.
  • 2025.09: One conference paper DeltaEnergy accepted to NeurIPS2025.
  • 2025.07: One extended paper Bayes-CAL accepted to IJCV.
  • 2025.03: One conference paper OODD accepted to CVPR2025.
  • 2025.01: One conference paper MaskST accepted to ICLR2025.
  • 2024.05: One conference paper CRoFT accepted to ICML2024.
  • 2024.04: One journal paper OoD-Control accepted to TPAMI.
  • 2024.02: One journal paper VLAD accepted to IJCV.
  • 2023.12: Two conference papers Bayes-CAL and SDL accepted to AAAI2023.

📝 Selected Publications

TPAMI
sym

InfoBound: A Provable Information-Bounds Inspired Framework for Both OoD Generalization and OoD Detection

Lin Zhu, Yifeng Yang, Zichao Nie, Yuan Gao, Jiarui Li, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye

This paper addresses real-world scenarios where covariate shifts and semantic shifts coexist, proposing a unified information-theoretic framework to improve both OoD generalization and OoD detection. By combining Mutual Information Minimization (MI-Min) and Conditional Entropy Maximization (CE-Max), the method mitigates the trade-off between the two tasks and achieves superior performance on multi-label image classification and object detection.

#

TPAMI
sym

CFSM: A Novel Causal Feature Selection Module for Two-Dimensional Out-of-Distribution Generalization

Lin Zhu, Weihan Yin, Yiyao Yang, Yifei Wu, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye

This paper identifies the limitations of existing causality-inspired OOD generalization methods, which mainly treat domain differences as confounders but may overlook complex spurious correlations in real-world data. To address this, the authors propose a modified causal intervention framework with a Causal Feature Selection Module (CFSM), which suppresses both domain-difference and spurious-correlation features, leading to improved two-dimensional OOD generalization performance.

#

ICLR2025
sym

Less is more: Masking elements in image condition features avoids content leakages in style transfer diffusion models

Lin Zhu, Xinbing Wang, Chenghu Zhou, Qinying Gu, Nanyang Ye

This paper addresses content leakage in style-reference-based text-to-image diffusion models, where reference images unintentionally transfer their content along with style. The authors propose a training-free masking strategy that drops selected image-feature elements from the style reference, showing that fewer but better-chosen conditions can improve content-style disentanglement and enhance style transfer performance.

#

IJCV
sym

Bayes-CAL: Robust Cross-Modal Alignment by Bayesian Approach for Few-Shot OoD Generalization

Lin Zhu*, Weihan Yin*, Fan Wu, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye ( *Equal contribution.)

This paper studies few-shot two-dimensional OoD generalization under both correlation shift and diversity shift, where fine-tuned large pre-trained models may overfit limited samples and generalize poorly to unseen classes. The authors propose Bayes-CAL, a Bayesian cross-modal image-text alignment method that fine-tunes text representations with domain-invariant regularization, achieving stronger and more stable OoD performance across image classification, object detection, and instance segmentation.

#

NeurIPS2025
sym

ΔEnergy: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization

Lin Zhu, Yifeng Yang, Xinbing Wang, Qinying Gu, Nanyang Ye

This paper proposes ΔEnergy, a new OOD score for vision-language models that better distinguishes semantic-shifted OOD samples by measuring energy changes after vision-language re-alignment. Based on ΔEnergy, the authors develop a unified fine-tuning framework that improves both OOD detection and covariate-shift OOD generalization, achieving substantial AUROC gains over recent methods.

#

ICML2024
sym

CRoFT: Robust Fine-Tuning with Concurrent Optimization for OOD Generalization and Open-Set OOD Detection

Lin Zhu, Yifeng Yang, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye

This paper studies how to fine-tune vision-language pre-trained models while preserving their ability to generalize under covariate shifts and detect semantic-shifted unseen classes. The authors propose a unified fine-tuning objective based on minimizing the gradient magnitude of energy scores, which theoretically promotes domain-consistent Hessians and empirically improves both OOD generalization and OOD detection.

#

IJCV
sym

Vision-Language Alignment Learning Under Affinity and Divergence Principles for Few-Shot Out-of-Distribution Generalization

Lin Zhu, Weihan Yin, Yiyao Yang, Fan Wu, Zhaoyu Zeng, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye

This paper proposes VLAD, a vision-language alignment framework for robust few-shot OOD generalization by using frozen CLIP language embeddings as invariant anchors and adapting visual features through lightweight adapters. By introducing affinity and divergence regularization, VLAD enhances class discrimination, suppresses non-causal features, and provides theoretical and empirical evidence for tighter OOD generalization bounds under distribution shifts.

📑 Working Papers

Adaptive Visual Evidence-Guided Decoding for Hallucination Reduction in LVLMs.

∆Energy++: Optimizing Energy Changes for Improved Robustness across Vision–Language Models.

ATT-COT: Prompting Vision-Language Models for OOD Generalization via Attention-Aware Chain-of-Thought.

🎖 Honors and Awards

  • National Scholarship for Postgraduate Students, 2025
  • NeurIPS Top Reviewer, 2025
  • Huatai Securities Technology Scholarship, 2024
  • Merit Student, Shanghai Jiao Tong University, 2024
  • Merit Student, Shanghai Jiao Tong University, 2023
  • Outstanding Graduates, ShanDong University, 2022
  • Excellent Student Cadre, Central South University, 2017

✔︎ Professional Activities

Reviewers