Projects
Research in MLLM understanding and diffusion models, 3D representation and agentic world generation, and adaptive reasoning and self-improving agents.
Research Projects
Building controllable image and video generation systems.
- Combined AST attention tuning, T³-S2S sketch tuning, and EGGen entity guidance to reduce missing objects and attribute leakage in multi-instance images. Under Review; TMLR 2026; ACM MM 2024.
- Used reward-aware sampling and diversity rewards in DRIFT to curb diversity collapse during image-model reinforcement fine-tuning while retaining alignment. NeurIPS 2026.
- Built XTalker and 3DXTalker with flow matching and separate speech, emotion, and motion controls for expressive 2D and 3D avatar animation. Under Review; arXiv 2026.
Representing editable 3D assets and generating coherent, observable, dynamic worlds.
- Developed P2Voxel pivot-voxel tokens and Hi-TOPS topology-aware parts to reduce mesh encoding cost and improve part decomposition. Under Review.
- Built StoryBlender with script-to-scene planning and spatial grounding to reduce identity and layout drift across editable 3D storyboard shots. ECCV 2026.
- Built Look-Before-Move with semantic observation targets and viewpoint search to choose visible subjects and feasible camera paths in dynamic 3D worlds. NeurIPS 2026.
- Used phase and role planning in Social Structure Matters to improve interaction timing and role consistency in two-person 3D world motion. Under Review.
Helping agents reason from evidence and adapt decisions across tasks.
- Developed Evolutionary Decoding with step-wise selection and block-wise mutation to escape confident errors in diffusion LLM mathematical reasoning. Under Review.
- Built MA-RAG to turn conflicting medical answers into retrieval queries, refining evidence across rounds to reduce hallucinations. ICML 2026.
- Developed S-ICQL with world-model-guided Q-learning and T2DA with language-decision alignment to adapt policies from trajectories or task text. ICLR 2026; NeurIPS 2025.
History Projects
Light-NAS is a distributed full-stack framework built with PyTorch and OpenMPI for efficient network design across classification, detection, 3D action recognition, latency prediction, MCUNet search, and quantization search.
- Designed entropy-based NAS algorithms for detection, action recognition, and quantization, leading to ICML 2022, NeurIPS 2022, and ICLR 2023 publications.
- Demonstrated practical acceleration of 3x for object detection, 1.68x for animation classfication, and 2.5x for car plate recognition.
- GitHub: alibaba/lightweight-neural-architecture-search
Learned compression methods use entropy models to approximate the distribution of compressible latents with CNNs, achieving performance comparable to traditional image codecs. In this line of work, we proposed the Interpolation Variable Rate model and Spatiotemporal Entropy approaches for image and video compression.
- 4 Tracks Winner of the Challenge on Learned Image Compression at CVPR 2019.
- Interpolation Variable Rate Image Compression accepted by ACM MM 2021.
- GitHub: tinyvision/IPCodec