Home
Fei Ma is a Research Scientist and Graduate Supervisor at Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(人工智能与数字经济广东省实验室(深圳)), also known as the Guangming Laboratory(光明实验室), where he leads the Multimedia (MM) Group under the general guidance of Prof. Qi Tian. His team focuses on multimodal content understanding and generation, with research spanning multimodal large language models, agents, AIGC, world models, and affective computing.
He received his Ph.D. in Information and Communication Engineering from Tsinghua University in 2022, and his B.Eng. in Communication Engineering from University of Electronic Science and Technology of China (UESTC) in 2017. He has published more than 40 papers in top-tier journals such as TPAMI, TMC, and TMLR, as well as at CCF-A conferences including ICML, NeurIPS, ICLR, CVPR, ACL, AAAI, and ACM MM. He has also filed or been granted over 40 Chinese invention patents. Prior to joining Guangming Laboratory, he worked at Huawei. This combined academic and industrial background drives his commitment to bridging the “last mile” between research breakthroughs and real-world deployment.
Research interns are welcome to apply year-round. If you are passionate about AI and eager to engage in cutting-edge research, please feel free to contact me by email.
长期招募研究型实习生。 如果您对人工智能充满热情,并渴望投身前沿研究,欢迎随时邮箱与我联系。
Recent News
-
[2026/07] We take 1st, 2nd, and 3rd place across three challenges at ACM MM 2026: EgoLink, AffectiveArt, and MPDD-AVG.
-
[2026/07] Four papers are accepted by ACM MM.
-
[2026/06] Three papers are accepted by ECCV.
-
[2026/04] Two papers are accepted by ICML.
-
[2026/04] Two papers are accepted by IJCAI.
-
[2026/03] We organize the 1st Workshop & Challenge on Neurophysiological Intelligence for Human-Aware Multimedia (NeuroMM) @ ACM MM 2026. Welcome to participate!
-
[2026/02] Three papers are accepted by CVPR.
-
[2026/01] Three papers are accepted by ICLR.
-
[2025/11] Our GMTalker project wins the Outstanding Scientific Research Achievement Innovation Award at the 27th China Hi-Tech Fair.
-
[2025/11] Two papers are accepted by AAAI.
-
[2025/09] Two papers are accepted by NeurIPS.
-
[2025/07] One paper is accepted by TPAMI.
-
[2025/07] Two papers are accepted by ACM MM.
-
[2025/05] One paper is accepted by ACL.
-
[2025/04] Three papers are accepted by IJCAI.
-
[2025/03] One paper is accepted by TPAMI. Congratulations to Hongwei!
Publications
- For more paper information, please refer to the Google Scholar page.
Journal Papers
-
[16] Y Wang, Y Zhu, Z Song, H Ma, Z Wang, Z Yu, F Ma, Q Tian. Uni-EmoAgent: A Unified Multi-Agent Framework for Affective Understanding and Art-Oriented Emotional Image Generation. IEEE Transactions on Affective Computing, 2026.
-
[15] J Ke, F Ma, Y Lin, X Wang, W Yang, K Liu, Z Yu, S Zhao, G Ding, Q Tian. RALoRA: Rank-Adaptive LoRA for Incomplete Multi-View Classification with Vision-Language Foundation Models. Information Fusion, 2026.
-
[14] S Ye, M Li, X Lin, J Zhang, W Xie, D Guo, F Ma, S Xia, Z Yu. CiMEU: Culture-Independent Multimodal Emotion Recognition via Disentangled Representation Learning. IEEE Transactions on Affective Computing, 2026.
-
[13] S Yang, W Yu, S Chen, F Ma, Z Liu, Q Li. A Terrain-Interactive Autonomous Switching Control Strategy for the Land-Air Bimodal Robot. IEEE Transactions on Intelligent Transportation Systems, 2026.
-
[12] Y Wang, H Yu, J Xu, F Ma, H Zhang, T Feng, Z Zhang, S Huang, D Sun, X Zhang. VFEM: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion. Transactions on Machine Learning Research, 2026.
-
[11] H Hu, Y Zhou, Q Wang, Y Zou, C Ma, J Si, J Liu, Z Yu, L Cui, F Ma, Q Tian. From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health. IEEE Transactions on Affective Computing, 2026.
-
[10] Y He, G Chen, F Yu, M Li, F Ma, G Zhou. Restoring neural radiance fields performance under adverse weather conditions. Engineering Applications of Artificial Intelligence, 2026.
-
[9] D Luo, H Xu, F Ma, N Zhang, L Wang, K Wang. Design and analysis of step apodized coupling surface grating for narrow linewidth distributed feedback lasers with highly efficient optical field modulation. Journal of Optics, 2026.
-
[8] R Shen, K Liu, B Zhang, W Yu, F Ma, Y Qu, Z Liu, Q Li. Trajectory Tracking Control of Fully Actuated Hexarotor UAVs With Adaptive Iterative Learning: From Theory to Application. IEEE Transactions on Industrial Electronics, 2025.
-
[7] S Chen, Z Wu, K Zhang, C Li, B Zhang, F Ma, F Yu, Q Li. Exploring embodied multimodal large models: Development, datasets, and future directions. Information Fusion, 2025.
-
[6] H Xue, X Luo, Z Hu, X Zhang, X Xiang, Y Dai, J Liu, Z Zhang, M Li, J Yang, F Ma, Z Wu, C Yang, Z Dai, F Yu. Human Motion Video Generation: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.
-
[5] H Hou, F Ma, Z Li, F Yu. VisualRWKV-HM: Enhancing Linear Visual-Language Models via Hybrid Mixing. Information Fusion, 2025.
-
[4] F Ma, Y Xie, Y Li, Y He, Y Zhang, H Ren, Z Liu, W Yao, F Ren, F Yu, S Ni. A Review of Human Emotion Synthesis Based on Generative Technology. IEEE Transactions on Affective Computing, 2025.
-
[3] H Ren, Y Zhou, J Zhu, X Lin, H Fu, Y Huang, Y Fang, F Ma, H Yu, B Cheng. Rethinking Efficient and Effective point-based Networks for Event Camera Classification and Regression. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.
-
[2] F Ma, Y Yuan, Y Xie, H Ren, I Liu, Y He, F Ren, F Yu, S Ni. Generative Technology for Human Emotion Recognition: A Scoping Review. Information Fusion, 2024.
-
[1] C Wang, H Yu, X Li, F Ma, X Wang, T Taleb, VCM Leung. Dependency-Aware Microservice Deployment for Edge Computing: A Deep Reinforcement Learning Approach with Network Representation. IEEE Transactions on Mobile Computing, 2024.
Conference Papers
-
[35] G Li, F Ma, H Xu, Q Tian. MEIR: Memory-Enhanced Incongruity-Aware Reasoning for Multimodal Sarcasm Detection and Explanation. ACM MM 2026.
-
[34] J Chen, J Kong, M Chen, B Zhang, S Zhang, H Zhao, R Huang, F Ma, Q Tian. UniGarment: Topology-Guided Texture Normalization for Simulation-Ready Garment Digitization. ACM MM 2026.
-
[33] Z Wang, Z Yu, Y Zhu, B Zhao, H Liang, T Wang, W Xia, J Zhang, Z Liu, H Ma, F Ma, Q Tian. AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition. ACM MM 2026.
-
[32] H Hu, L You, H Xu, Q Wang, F Yu, F Ma, Z Cheng, Z Lian, Y Zhou, L Cui. EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models. ACM MM 2026.
-
[31] S Liu, W Yu, B Zhang, S Chen, F Ma, Z Liu, Q Li. Disturbance-Aware Dynamical Trajectory Planning for Air-Land Bimodal Vehicles. IROS 2026.
-
[30] S Chen, Y Chen, Y Guan, Z Cheng, Z Zhang, S Qin, B Xia, J Li, W Yang, F Ma. Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding. ECCV 2026.
-
[29] C Guo, X Mo, Y Nie, F Ma, X Xu, C Long. TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding. ECCV 2026.
-
[28] D Zhu, R Hu, Z Yu, X Guo, M Wang, S Ye, F Ma, X Cao. MoMCE: Mixture of Modality and Cue Experts for Multimodal Deception Detection. ECCV 2026.
-
[27] Z Peng, E Yang, Y Cheng, H Yuan, F Ma, X Cao, L Shen. SCNS: Continual Personalization of Diffusion Models via Submodular Concept Neuron Selection. ICML 2026.
-
[26] H Ren, F Ma, X Lin, Y Fang, H Huang, Y Zhou, Y Huang, H Fu, Z Yang, Y Jiang, X Wu, B Cheng. Scalable Event Cloud Network for Event-based Classification. ICML 2026.
-
[25] X Li, D Yin, Y Xie, G Li, Z Cheng, F Ma. ViSA-Gait: Leveraging Vision Foundation Models for Semantic Anchored Gait Recognition. IJCAI 2026.
-
[24] M Yan, Y Shao, Y Pan, S Chen, H Pei, H Tang, F Ma, J Guo, N Sebe. Cross Domain Test Time Scaling: Scale knowledge and reasoning on cross domains. IJCAI 2026.
-
[23] J Chen, K Gao, Y Cui, M Sun, M Chen, S Wang, X Long, F Ma, Q Tian, R Huang, H Zhao. LottieGPT: Tokenizing Vector Animation for Autoregressive Generation. CVPR 2026.
-
[22] H Chang, X Xu, W Liu, J Wu, K Jiang, F Ma, Q Tian. TextOVSR: Text-Guided Real-World Opera Video Super-Resolution. CVPR 2026.
-
[21] C Guo, Y He, Y Nie, F Ma, X Xu, C Long. T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding. CVPR 2026.
-
[20] Z Chen, H Lin, Y Nie, F Ma, X Xu, F Yu, C Long. Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability. ICLR 2026.
-
[19] Y Huang, F Ma, Y Shao, J Guo, Z Yu, L Cui, Q Tian. Nuwa: Mending the Spatial Integrity Torn by VLM Token Pruning. ICLR 2026.
-
[18] Z Lian, L Sun, L Chen, H Chen, Z Cheng, F Zhang, Z Jia, Z Ma, F Ma, X Peng, J Tao. EmoPrefer: Can Large Language Models Understand Human Emotion Preferences? ICLR 2026.
-
[17] S Chen, T Zhao, Y Bin, F Ma, W Shao, Z Wang. D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies. AAAI 2026.
-
[16] J Jiang, Y Chen, P Chen, K Liu, J Zhou, Z Zhu, H Hu, F Ma, Q Tian, C Wu. A Principle-Driven Adaptive Policy for Group Cognitive Stimulation Dialogue for Elderly with Cognitive Impairment. AAAI 2026.
-
[15] Y Xie, R Min, Z Qin, F Ma, L Shen, F Yu, X Cao. RoMa: A Robust Model Watermarking Scheme for Protecting IP in Diffusion Models. NeurIPS 2025.
-
[14] Y Xie, M Li, S Li, X Li, G Chen, F Ma, F Yu, W Ding. Universal Visuo-Tactile Video Understanding for Embodied Interaction. NeurIPS 2025.
-
[13] H Wang, Q Li, L Chen, H Kang, F Ma, Y Jiang. HoloTrace: LLM-based Bidirectional Causal Knowledge Graph for Edge-Cloud Video Anomaly Detection. ACM MM 2025.
-
[12] Z Guo, Y Xie, W Xie, P Huang, C Wang, F Ma, F Yu. GaussianPU: Color Point Cloud Upsampling via 3D Gaussian Splatting. IROS 2025.
-
[11] Y Xie, B Ou, F Ma, Y Liu. Observation-Graph Interaction and Key-Detail Guidance for Vision and Language Navigation. IROS 2025.
-
[10] H Xue, Z Zhang, M Li, Z Dai, F Yu, F Ma, Z Wu. VideoHumanMIB: Unlocking Appearance Decoupling for Video Human Motion In-betweening. IJCAI 2025.
-
[9] W Feng, Y Zhu, R Zhang, C Wang, F Ma, X Wang, X Li. Active Multimodal Distillation for Few-shot Action Recognition. IJCAI 2025.
-
[8] G Chen, Y He, M Yu, F Yu, G Xu, F Ma, M Li, G Zhou. Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object Reconstruction. IJCAI 2025.
-
[7] C Zhang, J Peng, Z Wang, Y Lai, H Sun, H Chang, F Ma, W Yu. VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism. ACL 2025.
-
[6] Y Xie, T Feng, X Zhang, X Luo, Z Guo, W Yu, H Chang, F Ma, F Yu. PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis. AAAI 2025.
-
[5] X Xiang, Z Dai, H Xue, D Wang, M Li, Y Yue, F Ma, W Yu, H Chang, F Yu. ReMask-Animate: Refined Character Image Animation Using Mask-Guided Adapters. AAAI 2025.
-
[4] L Wang, S Shi, F Ma, F Yu, P Li, Y He. Subgraph Invariant Learning towards Large-scale Graph Node Classification. AAAI 2025.
-
[3] Z Zhong, Y He, P Li, F Yu, F Ma. A Language-Driven Navigation Strategy Integrating Semantic Maps and Large Language Models. IROS 2024.
-
[2] L Xiong, X Cheng, J Tan, X Wu, X Li, L Zhu, F Ma, M Li, H Xu, Z Hu. SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing. ACM MM 2024.
-
[1] X Luo, X Zhang, Y Xie, X Tong, W Yu, H Chang, F Ma, F Yu. CodeSwap: Symmetrically Face Swapping Based on Prior Codebook. ACM MM 2024.
Honors & Awards
- [2026/07] 1st Place in the Egocentric Language-Vision Interactive Network Knowledge Challenge (EgoLink) @ ACM MM 2026
- [2026/07] 2nd Place in the Challenge of Fine-Grained Emotion Understanding and Generation in Artistic Images (AffectiveArt) @ ACM MM 2026
- [2026/07] 3rd Place in the Challenge of Multimodal Personality-Aware Depression Detection via Audio-Visual Interview and Gait Analysis (MPDD-AVG) @ ACM MM 2026
- [2026/02] 光明实验室优秀团队
- [2026/02] 光明实验室优秀个人
- [2025/11] 中国国际高新技术成果交易会优秀科研成果创新奖
- [2024/12] 中国创新创业大赛创新挑战赛(宁波)解决方案优胜奖(赛道第一名)
- [2024/12] 深圳广告创意制作大赛AI创意生成类优秀奖
- [2024/09] 全国昇腾AI原生创新算子挑战赛(S2赛季)优秀奖(指导教师:马飞)
Membership
- 中国中文信息学会情感计算专委会委员
- 广东省图象图形学会情感计算专委会副秘书长
- 广东省青年科学家协会会员
- 深圳市哲学社会科学重点实验室主任助理
- 数字深圳联合创新中心专家委员会委员
- 深圳市光明社区科技委员
Invited Talks
- [2026/08/11] “迈向以人为中心的世界模型:初步探索” —— 鹏城实验室,深圳
- [2025/12/28] “基于大模型与AIGC的情感智能:从感知理解到具身共情” —— AIGC 2025(第三届人工智能生成内容国际会议暨大模型应用创新大会),杭州
- [2025/12/17] “生成式情感智能:从数字大脑到具身实体的跨越” —— 清华大学深圳国际研究生院,深圳
- [2025/10/28] “基于多模态大模型的媒体内容理解与生成” —— 中山市青联大讲堂,电子科技大学中山学院,中山
- [2025/10/17] “生成式情感智能:从多模态理解到具身交互” —— 安徽医科大学,合肥
- [2024/10/20] “多模态情感计算” —— 2024年第十三届华人心理学家学术研讨会,中山大学深圳校区,深圳
Academic Services
- Organizer/Chair:
ACM MM 2026 Workshop(NeuroMM 2026)
ACM MM 2026 Grand Challenge(NeuroMM 2026)
ACM MM 2026 Grand Challenge(MER 2026)
VALSE 2026 特色角(不止Paper,”攒局”指南)
- Journal Reviewer:
International Journal of Computer Vision
IEEE Transactions on Multimedia
IEEE Transactions on Circuits and Systems for Video Technology
IEEE Transactions on Affective Computing
IEEE Internet of Things Journal
Pattern Recognition
Neurocomputing
International Journal of Human-Computer Interaction
IEEE Robotics and Automation Letters
IEEE Signal Processing Letters
IEEE Transactions on Human-Machine Systems
Signal Processing: Image Communication
Software: Practice and Experience
Behaviour & Information Technology
Mobile Networks and Applications
Computers in Industry
Ocean Engineering
电子学报
- Conference Reviewer:
Conference on Neural Information Processing Systems(NeurIPS) (2025, 2023, 2022)
International Conference on Machine Learning(ICML) (2022)
International Conference on Learning Representations(ICLR)(2025,2024,2022)
Conference on Computer Vision and Pattern Recognition(CVPR) (2026,2025)
European Conference on Computer Vision(ECCV) (2026)
ACL Rolling Review (ARR) (2026)
AAAI Conference on Artificial Intelligence(AAAI) (2026)
ACM International Conference on Multimedia(ACM MM) (2025,2024)
IEEE International Conference on Multimedia and Expo(ICME) (2023,2022,2021)
ACM International Conference on Multimedia Retrieval(ICMR) (2026)
Winter Conference on Applications of Computer Vision(WACV) (2026)
International Conference on Image Processing(ICIP) (2021)
Teaching
- Generative Emotion Intelligence, Summer 2026, University of Electronic Science and Technology of China
- Data Thinking and Behavior, Fall 2025, Tsinghua University
- Machine Learning, Spring 2025, Shenzhen University
- Social Psychology and Behavioral Big Data, Spring 2023, Beijing Normal University at Zhuhai
- Information Theory, Fall 2018, Tsinghua University